A Kanban board with mandatory peer review gates and mechanical integrity enforcement, designed for concurrent AI agents.
btrain coordinates Claude Code, Codex, Gemini, and other AI agents working in the same repo. Lanes are WIP-limited work items moving through status columns (idle -> in-progress -> needs-review -> ready-for-pr -> pr-review -> ready-to-merge -> resolved), with feedback loops (changes-requested, repair-needed) that route work back to the writer or escalate to a human. Work is pulled, not pushed — agents claim lanes when capacity opens. File locks provide branch-level isolation without actual git branches, because AI agents share a single worktree. A watchdog auto-detects invalid state transitions, every active lane carries a structured delegation packet, and every handoff requires structured reviewer context (files changed, verification run, review asks) — essentially a PR description plus a worker contract built into the workflow.
The closest human equivalent: a team using a Kanban board where every card requires a PR approval before moving to "done", with automated integrity enforcement that a human board can't provide.
Human types "bth" in agent chat
|
v
btrain handoff <-- prints state + guidance
|
v
Agent follows instructions:
idle/resolved --> claim a task
in-progress --> continue working
needs-review --> review (if you're the reviewer)
ready-for-pr --> create/link a GitHub PR
pr-review --> wait for bot feedback, poll PR status
ready-to-merge --> merge the PR, then resolve
changes-requested --> fix findings, re-handoff
repair-needed --> fix workflow state
The source of truth is .claude/collab/HANDOFF_*.md (one per lane). Never edit these directly — always use the CLI.
PR flow is opt-in per repo:
[pr_flow]
enabled = true
base = "main"
required_bots = ["codex", "unblocked"]# Install
git clone https://github.com/codeslp/btrain.git && cd btrain && npm link
# Bootstrap a repo
btrain init /path/to/repo --agent "claude" --agent "codex" --agent "gemini"
# Launch the chat UI with all agents
btrain-chat all /path/to/repo
# Or launch a single agent
btrain-chat codex /path/to/repo
btrain-chat claude /path/to/repo
# Claim work
btrain handoff claim --lane a --task "Add auth middleware" \
--owner "claude" --reviewer "codex" --files "src/auth/"
# Check guidance
btrain handoff
# Hand off for review (poller auto-notifies the reviewer in #agents)
btrain handoff update --lane a --status needs-review --actor "claude" \
--preflight --changed "src/auth/index.ts" --verification "npm test" \
--why "Auth logic drifted" --review-ask "Check unauth flows"
# Keep this agent session attached to lane A and wake when its handoff changes.
# `bth wait` snapshots the current lane hash when --since is omitted.
bth wait --lane a --timeout 3600
# After a session restart, resume from a previously printed state hash.
bth wait --lane a --since <state-hash> --timeout 3600
# Reviewer approves
btrain handoff resolve --lane a --summary "Approved." --actor "codex"
# In repos with [pr_flow].enabled = true, local approval moves to ready-for-pr:
btrain pr create --lane a --bots all
btrain pr poll --lane a --apply
# If bots request changes, fix/push and request them again
btrain pr request-review --lane a --bots all
# When the PR is merged, poll once more to release locks and resolve the lane
btrain pr poll --lane a --apply
# Or requests changes
btrain handoff request-changes --lane a \
--summary "Need verification pass" --reason-code missing-verification \
--actor "codex"| Command | What it does |
|---|---|
btrain init <repo> |
Bootstrap handoff files, lanes, config, skills, dashboard, and agentchattr |
btrain handoff |
Print current state and what to do next |
btrain handoff wait |
Block on one lane's state hash, ignore unrelated lanes, then print the new canonical guidance; timeout exits 2 |
btrain handoff claim |
Claim a lane with task, owner, reviewer, file locks, and a delegation packet |
btrain handoff update |
Update status, delegation packet fields, and reviewer context |
btrain handoff resolve |
Approve local review; in PR-flow repos this advances to ready-for-pr |
btrain handoff request-changes |
Return review findings to the writer |
btrain pr create |
Push the branch, create a GitHub PR, link it to the lane, and request bot reviews |
btrain pr status |
Classify current PR bot review state |
btrain pr poll |
Fetch PR comments, classify feedback, and optionally update lane status |
btrain pr request-review |
Re-request configured bot reviews |
btrain status [--json] |
Show all lane states (JSON output for integrations) |
btrain repos [--json] |
List registered repos, including their enabled/disabled state |
btrain repos enable|disable|remove <name-or-path> |
Show/hide a repo globally or remove its registry record without deleting project files |
btrain repos prune |
Remove registry records for paths that no longer exist |
btrain doctor [--repair] [--skip-feedback] |
Health check; --repair fixes stale locks and workflow integrity |
btrain locks |
List active file locks across lanes |
btrain harness list |
List bundled and repo-local harness profiles plus the configured active profile |
btrain harness inspect |
Inspect one harness profile, its source path, and its probe-first metadata |
btrain startup |
Print a compact repo/bootstrap snapshot for a newly launched agent |
btrain hooks |
Install managed pre-commit + pre-push guards |
btrain override grant |
Human-confirmed override for blocked actions |
btrain hcleanup |
Trim handoff history |
Use the probe-first harness commands before changing workflow prompts or context wiring:
btrain harness list --repo .
btrain harness inspect --repo . --profile defaultbtrain harness list shows the configured active profile and every bundled or repo-local profile the loader can see. btrain harness inspect shows the selected profile's description, purpose, dispatch prompt, source path, and the registry/schema metadata behind it.
Repo-local overrides live under .btrain/harness/:
.btrain/harness/
├── registry.json
└── profiles/
└── <profile>.json
Use btrain startup --repo . for a compact first-turn packet: current branch and dirty summary, active agents/lanes, active harness profile, the most relevant lane focus, and the next commands to run.
Use btrain go --repo . when you need the broader bootstrap inventory of files and directories to read. startup is the quick orientation surface; go is the deeper file-reading checklist.
The dashboard is one global HUD for every repo in the BTrain registry. btrain init registers a repo; it does not create another dashboard process.
A local HUD at http://localhost:3333 with live lane status, hot seat indicators, and file locks:
btrain dashboard start # start and open the HUD
btrain dashboard status # print the process and URL
btrain dashboard open # reopen an existing HUD
btrain dashboard stop # stop the shared HUDThe first interactive btrain handoff claim in any registered repo starts and opens the HUD automatically. Claims from every other repo reuse that process and browser tab. The HUD reads btrain status --json, displays repo-qualified lanes, and lists stale registry entries separately instead of silently dropping them. The always-visible Manage repositories panel has a switch for every repo, so lanes can be hidden or restored directly from the dashboard. Disabled repos stay registered but are excluded from global status, sync operations, and dashboard lanes until re-enabled.
Use btrain repos prune to safely remove every registry entry whose directory no longer exists. btrain repos remove /path/to/repo only removes the registry record; it never deletes the repository or its files.
Set BTRAIN_DASHBOARD_AUTO_OPEN=false to disable claim-time startup, BTRAIN_DASHBOARD_DISABLED=true to disable all dashboard starts, or BTRAIN_DASHBOARD_PORT=<port> to choose the preferred port. If that port is occupied, btrain selects the next available port. Process state and logs are global at ~/.btrain/dashboard.json and ~/.btrain/dashboard.log (or under BRAIN_TRAIN_HOME). Existing repo-owned dashboards are stopped and migrated when that repo first starts the global HUD.
Features: all canonical handoff and PR states, compact status-colored lane indicators, per-repo on/off controls, hot-seat agent badges, expandable delegation details, repair-needed hazard styling, and a health endpoint for lifecycle monitoring.
A chat UI where Claude, Codex, and Gemini collaborate in real-time with live btrain integration.
Launch agents on any repo:
cd agentchattr && ./macos-linux/start_claude.sh # from the bootstrapped repo rootThen open http://127.0.0.1:8300 — the repo name shows in the header.
Features:
- Split-pane layout — lanes panel (left) + chat window (right)
- Mini lane pills — status-colored summary bar with dashboard animations
- Lane cards — status badge, LANE ID, W// R// agent meta, hot seat designation
- Live btrain state — 5s polling via
btrain status --json, WebSocket broadcasts - Auto-notify — poller auto-posts @mentions to
#agentson handoff transitions - Self-healing — content-hash dedup ensures notifications re-send if lost
- Auto-archive — lane messages archived when task changes
- Resizable/collapsible lanes panel with grip drag
- Click-to-switch — master-detail: click a lane card to switch the chat channel
Supported agents:
| Agent | CLI | Install | Auth |
|---|---|---|---|
| Claude | claude |
npm i -g @anthropic-ai/claude-code |
Subscription login |
| Codex | codex |
npm i -g @openai/codex |
Subscription login |
| Gemini | gemini |
npm i -g @google/gemini-cli |
Subscription login |
Also needed: tmux (brew install tmux) and ripgrep (brew install ripgrep) for Gemini.
Pre-flight readiness checks (btrain-chat status or GET /api/btrain/readiness) validate binary, auth, and working directory before launch.
Lane count = len(active agents) x per_agent:
1 agent, per_agent=3 --> lanes a, b, c
2 agents, per_agent=3 --> lanes a, b, c, d, e, f
3 agents, per_agent=3 --> lanes a, b, c, d, e, f, g, h, i
Config lives in .btrain/project.toml:
[project]
name = "your-repo"
[agents]
active = ["Claude", "GPT", "Gemini"]
writer_default = "Claude"
reviewer_default = "GPT"
[lanes]
enabled = true
per_agent = 3
[handoff]
# require_diff = false # Allow needs-review without a diff (docs-only repos)
[feedback]
# enabled = false # Opt out of feedback log scaffold and doctor warningsManage agents:
btrain agents set --repo . --agent "Claude" --agent "GPT" --lanes-per-agent 3
btrain agents add --repo . --agent "Gemini"Reviewer selection: --reviewer any-other picks any configured peer except the owner.
Every active lane may carry a structured Delegation Packet describing the contract for the work:
| Field | CLI flag | Purpose |
|---|---|---|
| Objective | --objective |
What the lane is trying to achieve |
| Deliverable | --deliverable |
Expected reviewable output or artifact |
| Constraints | --constraint (repeatable) |
Scope, safety, or workflow boundaries |
| Acceptance checks | --acceptance (repeatable) |
What must be true for the lane to count as done |
| Budget | --budget |
Effort or scope budget for the current pass |
| Done when | --done-when |
Concrete completion condition |
btrain handoff claim seeds these fields with useful defaults, and btrain handoff update can refine them as the lane is clarified or re-scoped.
| Status | Who acts | Description |
|---|---|---|
idle |
Anyone | Lane is free — claim a task |
in-progress |
Owner (writer) | Work underway |
needs-review |
Reviewer | Ready for peer review |
changes-requested |
Owner (writer) | Reviewer sent it back |
ready-for-pr |
Owner (writer) | Local review approved (PR-flow repos) — create or link the PR |
pr-review |
Owner (writer) | PR open — poll bot feedback until clear |
ready-to-merge |
Human/owner | Required bots clear — merge, then poll to resolve |
repair-needed |
Repair owner | Workflow integrity issue |
resolved |
Anyone | Review approved — lane recyclable |
Locks are retained through ready-for-pr, pr-review, and ready-to-merge,
and release when the PR merges. A PR that is closed without merging also
resolves the lane and releases its locks. Closure was previously routed to
repair-needed with the locks retained; the checks below caught that, and a
regression test now holds the corrected behavior in place.
Every needs-review handoff must include:
| Field | CLI flag | Purpose |
|---|---|---|
| Base | --base |
Branch or commit reference |
| Pre-flight review | --preflight |
What you checked before handoff |
| Files changed | --changed (repeatable) |
What files and why |
| Verification run | --verification (repeatable) |
Tests or checks that passed |
| Remaining gaps | --gap (repeatable) |
Unverified paths or known issues |
| Why this was done | --why |
Motivation for the change |
| Specific review asks | --review-ask (repeatable) |
What the reviewer should focus on |
btrain rejects needs-review transitions with placeholder context or empty diffs. Pass --no-diff to skip the diff gate for non-code changes (docs, config), or set require_diff = false under [handoff] in project.toml to disable it repo-wide.
changes-requested: spec-mismatch, regression-risk, missing-verification, security-risk, integration-breakage
repair-needed: invalid-handoff, unreviewed-push, lock-mismatch, ownership-conflict, state-conflict, invalid-transition, actor-mismatch, contradictory-state
btrain's coordination rules — who may move a lane, when files stay locked, and when locks release — are written down as a precise contract and checked by machines in two ways.
Design checking (pilot, not yet on main). A model checker
(TLA+/TLC) explores every reachable state of the intended workflow and
checks rules that must always hold: no two lanes ever lock the same files,
locks never outlive their lane, review always involves a second agent, and
no lane reaches the pull-request stage without a peer approval. The model
passes against a deliberately small configuration. The full rule list will
live in specs/tla/LaneLock.cfg, with the latest verified run cached beside
the model (both arrive with PR #35), so no counts are restated here to drift.
Code checking (available now). A test harness generates random workflow sequences (claims, reviews, PR outcomes, repairs), runs them against the real btrain code, and compares every step with the contract. A difference that is not already on the list of known open findings is shrunk to a minimal command sequence and reported with a seed and a trace that replay it exactly. Differences already on that list are counted by label instead, so they stay visible without failing the run that found them. This is how the rules stay enforced as the code changes, when the suite is run: a new regression against the contract fails a test instead of shipping quietly. It is opt-in locally and advisory in CI today, so run it before handing off.
npm installnpm run test:formal- Read the result. The command exits non-zero today, for two reasons that
mean different things:
- A held-open gap, working as designed. One check keeps the candidate
findings visible: differences between the code and the rules that are
still waiting on a decision about which side should change. It fails
while any candidate remains, so the gaps cannot be forgotten. Each is
described in
test/formal/README.md. - A test's stale copy of expected behavior, fixed in PR #34. One conformance check also failed because the test's copy of current behavior lagged a repair that had already landed. That was staleness in the test, not a regression; the fix is in PR #34 and the check is a regression signal again once it merges.
- Three earlier gaps have been fixed: a closed-but-unmerged PR left its files locked, a shortcut flag let a lane skip peer review, and dropping another lane's locks needed no audited sign-off. Each now has a passing regression test, and a return of any of them fails the suite.
- The regular
npm testsuite is unaffected; these checks are opt-in.
- A held-open gap, working as designed. One check keeps the candidate
findings visible: differences between the code and the rules that are
still waiting on a decision about which side should change. It fails
while any candidate remains, so the gaps cannot be forgotten. Each is
described in
- Optional settings:
BTRAIN_FORMAL_RUNS=<n>— random sequences per check (default 15, each 3 to 12 commands long)BTRAIN_FORMAL_SEED=<n>— replay a recorded run exactlyBTRAIN_FORMAL_TRACE_DIR=<path>— where failure traces are written
- No credentials, network, or AI provider is needed. Runs are repeatable from a seed.
The design check is not runnable from this branch yet. When the phase 1
work merges, specs/tla/ will carry the model, its configuration, the setup
instructions, and the latest cached result.
- Decide what kind of change you are making:
- Documentation or comments only (formatting, comments, docs that do
not change a workflow rule) — run the pin check,
python3 scripts/tla_pin.py --check, once that script exists (it lands with the phase 1 model in PR #35; until then there is nothing to check and the step is skipped). Each model records a fingerprint of the exact spec paragraphs it was built from; the check fails if you edited those paragraphs without deliberately updating the fingerprint, so a model can never silently drift from the prose it claims to follow. A prose edit that changes an authoritative rule inspecs/is a behavior change and follows the third path below, not this one. - Code changes that keep behavior the same (a refactor) — run
npm run test:formalto confirm the code still matches the rules. - Behavior changes — update the written rules first (in
specs/), then the model, then the code. The code follows the rules, never the reverse.
- Documentation or comments only (formatting, comments, docs that do
not change a workflow rule) — run the pin check,
- Work in your lane as usual, run the checks before handing off, and include in your review notes: which of the three kinds of change this is and why, which spec paragraphs it touches, the fingerprint-check verdict, the check results, and anything left unverified. The reviewer uses the first item to decide whether the checks you ran were the right ones.
- The reviewer confirms the rules say what was intended and that the checks ran against the exact change.
| Result | Meaning | What happens |
|---|---|---|
pass |
The design holds and the code matches the rules | Ready to review or merge |
stale_model |
The written rules changed but the model was not updated | Blocked until the model catches up |
counterexample |
The model checker found a sequence that breaks a safety rule | Blocked; the design or the rule needs fixing |
validation_mismatch |
The real code behaved differently from the rules | Blocked; fix the code or change the rules deliberately |
state_space_exhausted |
The model checker ran out of time or memory | Reported as a warning; a reviewer decides |
tool_unavailable |
A required tool or service was missing | Reported as an infrastructure problem, never as a pass or a failure |
policy_blocked |
An AI provider declined to perform a step | Reported separately; the step routes to an alternative or a human |
For the full policy, history, and current findings: the governing spec in
specs/, the harness and its findings ledger in test/formal/README.md,
and the tooling write-up in research/. One caveat while reading them: the
specs still describe the three fixed behaviors as outstanding drift. The
code and the regression tests are ahead of that prose, and correcting it is
tracked separately.
btrain does not measure token spend, and it does not manage context size.
Nothing in src/ reads usage data or enforces a context budget. This section
points at two external tools, and at the measurement that says which lever is
worth pulling.
Cache reads are roughly 70% of measured cost and scale with context size multiplied by turn count, so the levers are smaller payloads and fewer turns. A payload is billed as fresh input only on the turn it arrives and is re-read on every turn after, so shrinking what enters context helps twice.
npx ccusage@latest session— spend per session across claude, codex, and geminiast-grep— structural search, for when a regex would be fragile
Neither is a btrain dependency; both are run by hand. A context budget inside btrain is spec 020 workstream 3, which is not implemented.
See docs/token-tooling.md for when each tool helps, and spec 020 for the measurement, the date it was taken, and its scope — the figures are Claude-session only and they drift.
btrain scaffolds a .claude/collab/FEEDBACK_LOG.md during init and monitors it via btrain doctor.
feedback-triage skill — Triages user-reported feedback: assess category (MINOR/BUG/SPEC) and complexity, log to the feedback log, wait for user confirmation, then drive test-first resolution or route to speckit.
bug-fix skill — Test-first workflow for developer-found bugs: define the bug, write a failing reproduction test, fix against it, prove it passes.
btrain doctor warns on:
- Missing feedback log (on initialized repos)
- Stale entries (
newortriagedfor >7 days) - Entries missing required fields (Category, Status)
| Mechanism | Scope | Effect |
|---|---|---|
[feedback] enabled = false in project.toml |
Per-repo | Skips feedback log scaffold and all doctor warnings |
--skip-feedback on btrain doctor |
Per-run | Suppresses feedback warnings for one invocation |
--core-only on btrain init |
Per-repo | Skips bundled skills, local dashboard + agentchattr, feedback log, and feedback guidance in managed docs |
btrain init scaffolds these skills into .claude/skills/ and .agents/skills/ (skip with --core-only). Re-running init restores missing files but preserves existing skill content. Use btrain sync-skills --force --skill <name> to refresh an existing mirror from the bundled source; omit --skill to sync the whole bundle, and omit --force to preserve local edits while restoring missing files. Skill sync also refreshes the shared Unblocked helper required by the context-aware workflows.
| Skill | Purpose |
|---|---|
context-scout |
Risk-tier organizational context gathering with a standard btrain receipt |
feedback-triage |
Triage user-reported feedback into log, drive test-first resolution |
bug-fix |
Test-first bug investigation for developer-found bugs |
pre-handoff |
Quality gate before needs-review — catches placeholders, empty diffs |
reflect |
Post-failure reflection to turn incidents into prevention steps |
secure-by-default |
Trust-boundary check for auth, permissions, mutating endpoints |
integration-test-check |
Composition check when fixes span multiple components |
deploy-debug |
Classify deployment failures before debugging |
code-simplifier |
Simplify recent code and surface architecture-deepening opportunities |
test-writer |
Write or expand unit, integration, and component tests |
skill-creator |
Create or revise repo-local skills |
frontend-tokens |
Validate CSS custom property usage |
speckit-* |
Spec-driven development workflow (specify, clarify, plan, tasks, analyze, checklist, implement, taskstoissues) |
context-scout can use an existing zvec-grep index for one focused semantic discovery query. This route helps when an agent does not know the exact wording or file location. Native rg remains the required route for identifiers, paths, regular expressions, and exhaustive results.
btrain does not install zvec-grep, build an index, start its daemon, change global agent configuration, or authorize a remote Embedding provider. A user who wants the optional route must install and index the workspace explicitly. zvec-grep 0.2.x requires Node.js 22 or later.
npm install -g @zvec/zvec-grep
zg index --embedding local/potion-code-16m-v2 --hidden -g '!.agents/skills/**'btrain init adds .zvec-grep/ to .gitignore and scaffolds a read-only helper. The helper soft-skips (exit 0) when zg or a ready index is unavailable, when one zg call exceeds ZVEC_CONTEXT_TIMEOUT seconds (default 120), so a stale index under --freshness strict cannot block a task, or when it cannot create a temp file under TMPDIR. A search makes two bounded calls (readiness probe, then query), so its worst case is twice the bound.
.claude/scripts/zvec-context.sh status --root "$PWD"
# Low-risk orientation; may use the current index without refreshing it.
.claude/scripts/zvec-context.sh search \
"where lane transition authority is enforced" \
--root "$PWD" \
--limit 5
# Review, security, migration, and formal work must wait for a fresh index.
.claude/scripts/zvec-context.sh search \
"what validates a cached formal verdict against the current implementation" \
--root "$PWD" \
--freshness strict \
--glob 'specs/**' \
--glob 'scripts/**'The helper executes one hybrid query in direct mode. It never exposes index-management commands. Treat every ranked result as a lead and verify important claims with exact source reads.
btrain hooks or btrain init --hooks installs:
- pre-commit — blocks commits while a lane is waiting on review
- pre-push — blocks pushes while unresolved handoff state exists
Override: btrain override grant --action push --requested-by <agent> --confirmed-by <human> --reason "..."
btrain verifies which agent is speaking:
- Checks
BTRAIN_AGENTorBRAIN_TRAIN_AGENTenv var - Falls back to runtime hints +
[agents.runners]in config - Prints
agent check:line so agents can confirm identity
[agents.runners]
"Claude" = "claude -p"
"GPT" = "codex"
"Gemini" = "gemini"New and updated collaborator configurations infer these runner commands from common Claude, Codex/GPT, and Gemini names. Unknown agents receive an explicit notify runner so fallback behavior is visible in config. btrain doctor warns if an active agent has no runner mapping.
Use --lane when asking btrain to launch the agent whose turn it is:
btrain loop --lane b --max-rounds 1 --timeout 300The child process receives a pinned agent identity plus BTRAIN_LANE and BTRAIN_LANE_LOCKED=1. Its btrain and bth commands default to that lane and reject attempts to select a different lane. Dispatch traces are written under .btrain/harness/runs; inspect them with btrain harness trace list and btrain harness trace show <run-id>.
The original lane-less form remains available for single-lane repositories.
| Variable | Description |
|---|---|
BTRAIN_AGENT |
Pin the current agent identity |
BRAIN_TRAIN_AGENT |
Alternate variable for pinning agent identity |
BTRAIN_LANE |
Default lane for child btrain commands |
BTRAIN_LANE_LOCKED |
Reject commands that try to leave BTRAIN_LANE when set to 1 |
BRAIN_TRAIN_HOME |
Override the global btrain home directory |
HANDOFF_HISTORY_PATH |
Output file for the handoff history watcher |
btrain/
src/brain_train/
cli.mjs # CLI entry point
core.mjs # State machine, locks, events, watchdog
scripts/
serve-dashboard.js # Console HUD (port 3333)
agentchattr/
app.py # FastAPI server (port 8300)
wrapper.py # Agent CLI launcher + queue watcher
readiness.py # Pre-flight binary/auth/cwd checks
static/ # Chat UI (HTML/CSS/JS)
macos-linux/ # Launcher scripts per agent
specs/ # Design specs (004-007)
test/ # Node.js test suite
.btrain/ # Config, events, history, locks
.claude/collab/ # Handoff files (HANDOFF_A.md, etc.) + FEEDBACK_LOG.md
.claude/skills/ # Bundled + custom skills
.agents/skills/ # Bundled + custom skills for agents that read .agents
The agentchattr/ chat UI is built on agentchattr by @bcurts. btrain extends it with btrain lane integration, auto-notify handoff routing, split-pane lane cards, and multi-agent orchestration.
MIT