fix(aider): CONVENTIONS.md is a roster index, not 3.8 million characters of agents (#871)

Aider loads a conventions file into context and keeps it there for the whole
session — that is what the file is for, and the docs say to load it with
--read so prompt caching can hold it. This integration concatenated every
agent body into it. At 279 agents that is 3,816,372 characters, roughly a
million tokens. No model takes that. Anyone who ran

    ./scripts/install.sh --tool aider

got a CONVENTIONS.md that either blows the context window on the first turn
or bills for a million tokens trying.

CONVENTIONS.md is now the roster index it was described as: one entry per
agent with the name, the description, the division, and the path to the
agent file. 96,823 characters, down from 3.8 million. The header explains
how to pull a single agent's full instructions into the session:

    /read-only /path/to/agency-agents/engineering/engineering-frontend-developer.md

Naming an agent in a prompt still works the way it did — the description is
what the model needed for that, and it is still there.

test-convert-outputs.sh now holds the index to being an index: it fails if
CONVENTIONS.md grows past 250,000 characters, if it does not list exactly
one path per roster agent, or if any path it prints does not resolve. The
existing round-trip check on the accumulated file still covers the names and
descriptions.
This commit is contained in:
Hotragn Pettugani
2026-09-20 18:35:10 -05:00
committed by GitHub
parent ad9264e309
commit d3a3f573e3
6 changed files with 89 additions and 22 deletions
+25
View File
@@ -353,6 +353,31 @@ if split_bad:
else:
ok(f"openclaw: all {N} agents keep every source fenced block whole in one output file")
# --- Layer A (context budget): the Aider index has to stay an index ----------
# Aider keeps a conventions file in context for the whole session. Inlining the
# agent bodies made CONVENTIONS.md 3.8 million characters, which no model will
# take, so it carries one index entry per agent instead: description plus the
# path to the real file. Two things have to hold for that to be worth anything —
# the file stays small enough to load, and every path it prints resolves.
AIDER_INDEX_CEILING = 250_000
aider_index = os.path.join(OUT, "aider", "CONVENTIONS.md")
if os.path.isfile(aider_index):
text = open(aider_index, encoding="utf-8").read()
if len(text) > AIDER_INDEX_CEILING:
bad(f"aider: CONVENTIONS.md is {len(text):,} characters — it is loaded into "
f"every request, so it has to stay an index, not the agents themselves")
paths = re.findall(r"^Full instructions: (.+)$", text, re.M)
dangling = sorted({p for p in paths if not os.path.isfile(os.path.join(R, p))})
if len(paths) != N:
bad(f"aider: CONVENTIONS.md points at {len(paths)} agent files, roster has {N}")
elif dangling:
for d in dangling[:3]:
bad(f"aider: CONVENTIONS.md points at a file that does not exist: {d}")
if len(dangling) > 3:
bad(f"aider: ...and {len(dangling)-3} more dangling paths")
elif len(text) <= AIDER_INDEX_CEILING:
ok(f"aider: index is {len(text):,} characters and all {N} agent paths resolve")
# --- Layer A (app-facing): every SOURCE frontmatter strict-parsed above -------
for m in src_bad[:5]: bad(m)
if len(src_bad) > 5: bad(f"...and {len(src_bad)-5} more source frontmatter problems")