test(install): add a regression suite for install.sh + CI on Linux and macOS (#772)

* test(install): add a regression suite for install.sh + CI on Linux and macOS

install.sh is the largest script in the repo and has no tests. Every install
bug so far has been a silent one — agents copied to the wrong directory, a
path with a space split in two, a filter that installed everything — and the
only signal was a user noticing later.

scripts/test-install.sh pins the installer's observable contract:

  * destinations: default $HOME/.claude/agents, --path override, tool env var
    override, and --path winning over the env var
  * selection: --division, --agent, --agents-file (comments/blank lines)
  * --dry-run writes nothing; unknown --tool exits non-zero
  * --link produces symlinks; a second run installs the same set, not dupes
  * a destination containing spaces stays one directory

Expected counts are derived from divisions.json + lib.sh at runtime, so the
suite doesn't need updating when agents are added. Every case runs with HOME
pointed at a throwaway sandbox, so a broken default path can never write into
the real config. bash 3.2 + BSD userland, no new dependencies.

Verified it fails on the regressions it claims to catch: unquoting install_file
fails only the spaces case, neutering slug_allowed fails the four selection
cases, un-short-circuiting --dry-run fails the dry-run case, and ignoring the
env var in resolve_dest fails the env-override case.

CI runs it on ubuntu-latest and macos-latest (macOS ships bash 3.2, Linux
ships bash 5) plus bash -n over every script in scripts/.

* test(install): pin the parallel worker argument regression (#755)

Review feedback: the existing "paths with spaces" case selects a single tool,
so it stays on the serial path and never reaches the worker spawn where #755's
bug lives. Adds a case that does.

  --tool claude-code,copilot --parallel --jobs 1 --agents-file <spaced path>
  --path "<home>/My [Agents]/dest dir"

with a serial control immediately before it (same two tools, same spaced and
globbed --path, no --parallel) so a failure is attributable to the worker
hand-off rather than to the selection filter.

Marked xfail rather than a hard assertion: it fails on main today and passes
with #755 applied, and encoding a known-broken case as a hard failure would
turn CI red for reasons unrelated to whatever PR is being reviewed. xfail
never fails the suite; when the case starts passing it prints a note to
promote it to assert_eq (one-word edit). Measured on macOS bash 3.2.57:
main -> 25 passed / 1 xfail, #755 applied -> 26 passed / 0 failed, both
deterministic over repeated runs.

Note on --jobs 1: workers are still spawned through the same xargs/sh
hand-off, so argument propagation is exercised in full. Serializing them
keeps a second, unrelated defect out of this case — with two workers running
concurrently against one shared --path, the parent exits non-zero on ~3 runs
in 5 once the workers actually copy anything (one worker's cp fails with
ENOENT on the shared destination). That race is invisible on main only
because the workers currently install nothing at all; --jobs 1 or per-tool
destinations are clean. Reported in the PR discussion.
This commit is contained in:
Sergio Romero
2026-09-02 20:48:48 -05:00
committed by GitHub
parent 3febe026c1
commit 1e2d1b940e
3 changed files with 337 additions and 0 deletions
+30
View File
@@ -0,0 +1,30 @@
name: Test Installer
# No path filter on purpose: the installer's contract can break from the other
# side too — a renamed division, a file that loses its frontmatter, a change to
# lib.sh — so these run on every PR.
on:
pull_request:
push:
branches: [main]
jobs:
test-install:
name: install.sh behavior (${{ matrix.os }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
# macOS ships bash 3.2, Linux ships bash 5 — the scripts must pass on both.
os: [ubuntu-latest, macos-latest]
steps:
- uses: actions/checkout@v4
- name: Shell syntax
run: |
for f in scripts/*.sh; do bash -n "$f"; done
- name: Run installer tests
run: |
chmod +x scripts/test-install.sh scripts/install.sh
./scripts/test-install.sh
+5
View File
@@ -240,6 +240,11 @@ Want agency-agents to install into a new tool (a CLI, editor, or agent runtime)?
4. **`.gitignore`** — add a rule for your tool's generated output under `integrations/<tool>/`. **This step is required and easy to miss.** Converted agent/skill files are generated locally by `convert.sh` and are **never committed** (see "Things we'll always close" below) — only `integrations/<tool>/README.md` is tracked. Match an existing per-tool entry.
5. **`integrations/<tool>/README.md`** — a short doc for the integration (every tool has one; it's the only committed file in the tool's directory).
6. **Run `./scripts/check-tools.sh`** — it must pass. It cross-checks `tools.json` against `install.sh` and `convert.sh` and flags anything missing.
7. **Run `./scripts/test-install.sh`** — it must pass. It installs into throwaway
sandboxes (never your real `$HOME`) and pins the installer's observable
contract: where files land, that `--path` beats the tool's env var, that
`--division` / `--agent` / `--agents-file` filter, that `--dry-run` writes
nothing, and that paths with spaces survive. CI runs it on Linux and macOS.
If your PR commits the converted output (the generated `integrations/<tool>/*` files), CI and review will ask you to remove it and add the `.gitignore` rule instead.
+302
View File
@@ -0,0 +1,302 @@
#!/usr/bin/env bash
#
# test-install.sh — regression tests for scripts/install.sh.
#
# install.sh is the largest script in the repo and every install bug so far has
# been a silent one: agents land in the wrong directory, a path with a space is
# split into two, a filter installs everything. These tests pin the observable
# contract — where files land and how many — so those regressions fail loudly.
#
# Design constraints (same as the rest of scripts/):
# * bash 3.2 + BSD userland, no jq, no GNU-only flags.
# * Never touches the real $HOME. Every case runs with HOME set to a fresh
# sandbox, so a broken default path writes into the sandbox, not your config.
# * Only exercises the two source tools (claude-code, copilot) so no case
# depends on convert.sh output being present or fresh.
#
# Usage: ./scripts/test-install.sh [-v]
# -v echo the installer's own output for failing cases
set -uo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
INSTALL="$SCRIPT_DIR/install.sh"
# shellcheck source=scripts/lib.sh
. "$SCRIPT_DIR/lib.sh"
VERBOSE=false
[[ "${1:-}" == "-v" ]] && VERBOSE=true
passed=0
failed=0
xfailed=0
SANDBOX_ROOT="$(mktemp -d "${TMPDIR:-/tmp}/agency-install-tests.XXXXXX")"
trap 'rm -rf "$SANDBOX_ROOT"' EXIT
pass() { printf ' ok %s\n' "$1"; passed=$((passed + 1)); }
fail() {
printf ' FAIL %s\n' "$1"
[[ -n "${2:-}" ]] && printf ' %s\n' "$2"
failed=$((failed + 1))
}
# assert_eq <expected> <actual> <label>
assert_eq() {
if [[ "$1" == "$2" ]]; then pass "$3"; else fail "$3" "expected '$1', got '$2'"; fi
}
# xfail_eq <expected> <actual> <label> <tracking> — a case that is known to fail
# until a specific fix lands. It never turns the suite red: a mismatch is the
# documented status quo today, and a match means the fix landed and the case
# should be promoted to a plain assert_eq (one-word edit).
xfail_eq() {
if [[ "$1" == "$2" ]]; then
pass "$3"
printf ' ^ %s appears to have landed — promote this case to assert_eq\n' "$4"
else
printf ' xfail %s\n' "$3"
printf ' expected '\''%s'\'', got '\''%s'\'' — fixed by %s\n' "$1" "$2" "$4"
xfailed=$((xfailed + 1))
fi
}
# sandbox <name> — fresh HOME for one case; echoes its path.
sandbox() {
local d="$SANDBOX_ROOT/$1"
rm -rf "$d"; mkdir -p "$d"
printf '%s' "$d"
}
# run_install <home> [args...] — run the installer with an isolated HOME.
# Captures stdout+stderr in RUN_OUT and the exit status in RUN_STATUS.
run_install() {
local home="$1"; shift
RUN_OUT="$(HOME="$home" "$INSTALL" --no-interactive "$@" 2>&1)"
RUN_STATUS=$?
$VERBOSE && printf '%s\n' "$RUN_OUT"
return 0
}
# count_md <dir> — .md files directly in <dir> (0 when the dir does not exist).
count_md() {
[[ -d "$1" ]] || { printf '0'; return; }
find "$1" -maxdepth 1 -name '*.md' -type f | wc -l | tr -d ' '
}
# ---------------------------------------------------------------------------
# Expected values, derived from the repo the same way install.sh derives them
# (divisions.json -> directories -> files with frontmatter), never hardcoded.
# ---------------------------------------------------------------------------
divisions_from_json() {
awk '/"divisions"[[:space:]]*:[[:space:]]*\{/{f=1; next} f' "$REPO_ROOT/divisions.json" \
| grep -oE '^[[:space:]]*"[a-z0-9-]+"[[:space:]]*:' \
| sed -E 's/[[:space:]]*"([a-z0-9-]+)"[[:space:]]*:/\1/'
}
ALL_DIVISIONS=()
while IFS= read -r _d; do [[ -n "$_d" ]] && ALL_DIVISIONS+=("$_d"); done < <(divisions_from_json)
agent_files_in() {
local d="$REPO_ROOT/$1" f
[[ -d "$d" ]] || return 0
while IFS= read -r f; do is_agent_file "$f" && printf '%s\n' "$f"; done \
< <(find "$d" -name "*.md" -type f | sort)
}
TOTAL_AGENTS=0
for _div in "${ALL_DIVISIONS[@]}"; do
TOTAL_AGENTS=$(( TOTAL_AGENTS + $(agent_files_in "$_div" | wc -l | tr -d ' ') ))
done
ENG_AGENTS=$(agent_files_in engineering | wc -l | tr -d ' ')
# `awk NR==1` rather than `head -1`: head exits at the first line and the
# still-writing function dies on SIGPIPE, which prints a spurious error.
FIRST_ENG_FILE="$(agent_files_in engineering | awk 'NR==1')"
FIRST_ENG_SLUG="$(agent_slug "$FIRST_ENG_FILE")"
echo "Testing $INSTALL"
echo " repo: $REPO_ROOT"
echo " ${#ALL_DIVISIONS[@]} divisions, $TOTAL_AGENTS agents (engineering: $ENG_AGENTS)"
echo ""
# ---------------------------------------------------------------------------
# 1. Help and listings
# ---------------------------------------------------------------------------
echo "help + listings"
home="$(sandbox help)"
RUN_OUT="$(HOME="$home" "$INSTALL" --help 2>&1)"; RUN_STATUS=$?
assert_eq 0 "$RUN_STATUS" "--help exits 0"
case "$RUN_OUT" in *"Usage:"*) pass "--help prints usage" ;; *) fail "--help prints usage" ;; esac
home="$(sandbox list-teams)"
run_install "$home" --list teams
assert_eq 0 "$RUN_STATUS" "--list teams exits 0"
missing=""
for _div in "${ALL_DIVISIONS[@]}"; do
case "$RUN_OUT" in *"$_div"*) ;; *) missing="$missing $_div" ;; esac
done
assert_eq "" "$missing" "--list teams names every division in divisions.json"
home="$(sandbox list-agents)"
run_install "$home" --list agents
listed=$(printf '%s\n' "$RUN_OUT" | grep -c "$FIRST_ENG_SLUG")
[[ "$listed" -ge 1 ]] && pass "--list agents includes $FIRST_ENG_SLUG" \
|| fail "--list agents includes $FIRST_ENG_SLUG"
# ---------------------------------------------------------------------------
# 2. --dry-run writes nothing
# ---------------------------------------------------------------------------
echo ""
echo "dry-run"
home="$(sandbox dry-run)"
run_install "$home" --tool claude-code --dry-run
assert_eq 0 "$RUN_STATUS" "--dry-run exits 0"
assert_eq 0 "$(find "$home" -type f | wc -l | tr -d ' ')" "--dry-run creates no files"
# ---------------------------------------------------------------------------
# 3. Default destination + --path override
# ---------------------------------------------------------------------------
echo ""
echo "destinations"
home="$(sandbox default-dest)"
run_install "$home" --tool claude-code
assert_eq "$TOTAL_AGENTS" "$(count_md "$home/.claude/agents")" \
"claude-code installs every agent to \$HOME/.claude/agents"
assert_eq 0 "$(count_md "$home/.claude")" "claude-code writes nothing into the config root"
home="$(sandbox path-override)"
dest="$home/custom-dir"
run_install "$home" --tool claude-code --path "$dest"
assert_eq "$TOTAL_AGENTS" "$(count_md "$dest")" "--path overrides the default destination"
assert_eq 0 "$(count_md "$home/.claude/agents")" "--path leaves the default destination empty"
# Env var override, and --path winning over it. COPILOT_AGENT_DIR is used here
# because it unambiguously names the agents directory itself.
home="$(sandbox env-override)"
dest="$home/from-env"
RUN_OUT="$(HOME="$home" COPILOT_AGENT_DIR="$dest" "$INSTALL" --no-interactive --tool copilot 2>&1)"
assert_eq "$TOTAL_AGENTS" "$(count_md "$dest")" "COPILOT_AGENT_DIR overrides the default destination"
home="$(sandbox env-vs-path)"
RUN_OUT="$(HOME="$home" COPILOT_AGENT_DIR="$home/from-env" "$INSTALL" --no-interactive \
--tool copilot --path "$home/from-flag" 2>&1)"
assert_eq "$TOTAL_AGENTS" "$(count_md "$home/from-flag")" "--path wins over the env var"
assert_eq 0 "$(count_md "$home/from-env")" "env var destination is unused when --path is given"
# ---------------------------------------------------------------------------
# 4. Paths with spaces (regression: word-splitting in the install loop)
# ---------------------------------------------------------------------------
echo ""
echo "paths with spaces"
home="$(sandbox 'spaces')"
dest="$home/My Agents/claude code"
run_install "$home" --tool claude-code --path "$dest"
assert_eq "$TOTAL_AGENTS" "$(count_md "$dest")" "installs into a path containing spaces"
assert_eq 0 "$(find "$home" -maxdepth 1 -name 'My' -o -maxdepth 1 -name 'Agents' | wc -l | tr -d ' ')" \
"a spaced path is not split into separate directories"
# ---------------------------------------------------------------------------
# 4b. Parallel workers get their arguments intact (PR #755)
#
# --parallel hands the parent's selection state to child workers. On main that
# happens through a command-shaped string expanded unquoted, so a --path or
# --agents-file containing whitespace or glob characters is word-split and
# pathname-expanded on the way in. The install then writes nothing while still
# reporting "Done! Installed 2 tool(s)" and exiting 0 — a silent miss, which is
# why the exit-status assertion below cannot catch it on its own and the file
# count is what actually pins the regression.
#
# Two tools are required: a single tool stays on the serial path and never
# reaches the worker spawn.
#
# --jobs 1 is deliberate. Workers are still spawned through the same xargs/sh
# hand-off, so argument propagation — the thing under test — is exercised in
# full; serializing them just keeps a second, unrelated defect out of this case.
# With two workers running concurrently against one shared --path, the parent
# exits non-zero on roughly 3 runs in 5 (measured on macOS, bash 3.2) once the
# workers actually copy anything. That race is invisible on main only because
# the workers currently install nothing at all. See the PR discussion.
# ---------------------------------------------------------------------------
echo ""
echo "parallel workers"
home="$(sandbox parallel-serial-control)"
dest="$home/My [Agents]/dest dir"
list="$home/my agents list.txt"
{ echo "# same selection as the parallel case below"; echo "$FIRST_ENG_SLUG"; } > "$list"
run_install "$home" --tool claude-code,copilot --no-convert --agents-file "$list" --path "$dest"
assert_eq 0 "$RUN_STATUS" "serial control: two tools, spaced/globbed --path, exits 0"
assert_eq 1 "$(count_md "$dest")" "serial control: installs exactly the one selected agent"
home="$(sandbox parallel)"
dest="$home/My [Agents]/dest dir"
list="$home/my agents list.txt"
{ echo "# one agent, listed in a file whose own path has spaces"; echo "$FIRST_ENG_SLUG"; } > "$list"
run_install "$home" --tool claude-code,copilot --parallel --jobs 1 --no-convert --agents-file "$list" --path "$dest"
assert_eq 0 "$RUN_STATUS" "--parallel with a spaced/globbed --path exits 0"
xfail_eq 1 "$(count_md "$dest")" \
"--parallel installs exactly the one selected agent (spaced --path + --agents-file)" "PR #755"
# ---------------------------------------------------------------------------
# 5. Selection filters
# ---------------------------------------------------------------------------
echo ""
echo "selection"
home="$(sandbox division)"
dest="$home/dest"
run_install "$home" --tool claude-code --division engineering --path "$dest"
assert_eq "$ENG_AGENTS" "$(count_md "$dest")" "--division installs only that division"
home="$(sandbox agent)"
dest="$home/dest"
run_install "$home" --tool claude-code --agent "$FIRST_ENG_SLUG" --path "$dest"
assert_eq 1 "$(count_md "$dest")" "--agent installs exactly one agent"
home="$(sandbox agents-file)"
dest="$home/dest"
list="$home/agents.txt"
{ echo "# comment line"; echo ""; echo "$FIRST_ENG_SLUG"; } > "$list"
run_install "$home" --tool claude-code --agents-file "$list" --path "$dest"
assert_eq 1 "$(count_md "$dest")" "--agents-file skips comments and blank lines"
home="$(sandbox unknown-tool)"
run_install "$home" --tool definitely-not-a-tool
[[ "$RUN_STATUS" -ne 0 ]] && pass "unknown --tool exits non-zero" \
|| fail "unknown --tool exits non-zero" "exited 0"
# ---------------------------------------------------------------------------
# 6. --link and idempotency
# ---------------------------------------------------------------------------
echo ""
echo "link + repeat runs"
home="$(sandbox link)"
dest="$home/dest"
run_install "$home" --tool claude-code --link --division engineering --path "$dest"
links=$(find "$dest" -maxdepth 1 -type l | wc -l | tr -d ' ')
assert_eq "$ENG_AGENTS" "$links" "--link creates symlinks, not copies"
home="$(sandbox idempotent)"
dest="$home/dest"
run_install "$home" --tool claude-code --division engineering --path "$dest"
first=$(count_md "$dest")
run_install "$home" --tool claude-code --division engineering --path "$dest"
assert_eq "$first" "$(count_md "$dest")" "re-running installs the same set, not duplicates"
# ---------------------------------------------------------------------------
echo ""
if [[ $xfailed -gt 0 ]]; then
echo "Results: $passed passed, $failed failed, $xfailed known-broken (xfail)."
else
echo "Results: $passed passed, $failed failed."
fi
if [[ $failed -gt 0 ]]; then
echo "FAILED"
exit 1
fi
echo "PASSED"