Skip to content

Repository files navigation

btrain

A Kanban board with mandatory peer review gates and mechanical integrity enforcement, designed for concurrent AI agents.

btrain coordinates Claude Code, Codex, Gemini, and other AI agents working in the same repo. Lanes are WIP-limited work items moving through status columns (idle -> in-progress -> needs-review -> ready-for-pr -> pr-review -> ready-to-merge -> resolved), with feedback loops (changes-requested, repair-needed) that route work back to the writer or escalate to a human. Work is pulled, not pushed — agents claim lanes when capacity opens. File locks provide branch-level isolation without actual git branches, because AI agents share a single worktree. A watchdog auto-detects invalid state transitions, every active lane carries a structured delegation packet, and every handoff requires structured reviewer context (files changed, verification run, review asks) — essentially a PR description plus a worker contract built into the workflow.

The closest human equivalent: a team using a Kanban board where every card requires a PR approval before moving to "done", with automated integrity enforcement that a human board can't provide.

Node.js 18+ Zero Dependencies License: MIT


Overview

Human types "bth" in agent chat
       |
       v
btrain handoff   <-- prints state + guidance
       |
       v
Agent follows instructions:
  idle/resolved   --> claim a task
  in-progress     --> continue working
  needs-review    --> review (if you're the reviewer)
  ready-for-pr     --> create/link a GitHub PR
  pr-review        --> wait for bot feedback, poll PR status
  ready-to-merge   --> merge the PR, then resolve
  changes-requested --> fix findings, re-handoff
  repair-needed   --> fix workflow state

The source of truth is .claude/collab/HANDOFF_*.md (one per lane). Never edit these directly — always use the CLI.

PR flow is opt-in per repo:

[pr_flow]
enabled = true
base = "main"
required_bots = ["codex", "unblocked"]

Quick Start

# Install
git clone https://github.com/codeslp/btrain.git && cd btrain && npm link

# Bootstrap a repo
btrain init /path/to/repo --agent "claude" --agent "codex" --agent "gemini"

# Launch the chat UI with all agents
btrain-chat all /path/to/repo

# Or launch a single agent
btrain-chat codex /path/to/repo
btrain-chat claude /path/to/repo

# Claim work
btrain handoff claim --lane a --task "Add auth middleware" \
  --owner "claude" --reviewer "codex" --files "src/auth/"

# Check guidance
btrain handoff

# Hand off for review (poller auto-notifies the reviewer in #agents)
btrain handoff update --lane a --status needs-review --actor "claude" \
  --preflight --changed "src/auth/index.ts" --verification "npm test" \
  --why "Auth logic drifted" --review-ask "Check unauth flows"

# Keep this agent session attached to lane A and wake when its handoff changes.
# `bth wait` snapshots the current lane hash when --since is omitted.
bth wait --lane a --timeout 3600

# After a session restart, resume from a previously printed state hash.
bth wait --lane a --since <state-hash> --timeout 3600

# Reviewer approves
btrain handoff resolve --lane a --summary "Approved." --actor "codex"

# In repos with [pr_flow].enabled = true, local approval moves to ready-for-pr:
btrain pr create --lane a --bots all
btrain pr poll --lane a --apply

# If bots request changes, fix/push and request them again
btrain pr request-review --lane a --bots all

# When the PR is merged, poll once more to release locks and resolve the lane
btrain pr poll --lane a --apply

# Or requests changes
btrain handoff request-changes --lane a \
  --summary "Need verification pass" --reason-code missing-verification \
  --actor "codex"

What's Included

btrain CLI

Command What it does
btrain init <repo> Bootstrap handoff files, lanes, config, skills, dashboard, and agentchattr
btrain handoff Print current state and what to do next
btrain handoff wait Block on one lane's state hash, ignore unrelated lanes, then print the new canonical guidance; timeout exits 2
btrain handoff claim Claim a lane with task, owner, reviewer, file locks, and a delegation packet
btrain handoff update Update status, delegation packet fields, and reviewer context
btrain handoff resolve Approve local review; in PR-flow repos this advances to ready-for-pr
btrain handoff request-changes Return review findings to the writer
btrain pr create Push the branch, create a GitHub PR, link it to the lane, and request bot reviews
btrain pr status Classify current PR bot review state
btrain pr poll Fetch PR comments, classify feedback, and optionally update lane status
btrain pr request-review Re-request configured bot reviews
btrain status [--json] Show all lane states (JSON output for integrations)
btrain repos [--json] List registered repos, including their enabled/disabled state
btrain repos enable|disable|remove <name-or-path> Show/hide a repo globally or remove its registry record without deleting project files
btrain repos prune Remove registry records for paths that no longer exist
btrain doctor [--repair] [--skip-feedback] Health check; --repair fixes stale locks and workflow integrity
btrain locks List active file locks across lanes
btrain harness list List bundled and repo-local harness profiles plus the configured active profile
btrain harness inspect Inspect one harness profile, its source path, and its probe-first metadata
btrain startup Print a compact repo/bootstrap snapshot for a newly launched agent
btrain hooks Install managed pre-commit + pre-push guards
btrain override grant Human-confirmed override for blocked actions
btrain hcleanup Trim handoff history

Harness Discovery

Use the probe-first harness commands before changing workflow prompts or context wiring:

btrain harness list --repo .
btrain harness inspect --repo . --profile default

btrain harness list shows the configured active profile and every bundled or repo-local profile the loader can see. btrain harness inspect shows the selected profile's description, purpose, dispatch prompt, source path, and the registry/schema metadata behind it.

Repo-local overrides live under .btrain/harness/:

.btrain/harness/
├── registry.json
└── profiles/
    └── <profile>.json

Startup vs Go

Use btrain startup --repo . for a compact first-turn packet: current branch and dirty summary, active agents/lanes, active harness profile, the most relevant lane focus, and the next commands to run.

Use btrain go --repo . when you need the broader bootstrap inventory of files and directories to read. startup is the quick orientation surface; go is the deeper file-reading checklist.

Console Dashboard

The dashboard is one global HUD for every repo in the BTrain registry. btrain init registers a repo; it does not create another dashboard process.

A local HUD at http://localhost:3333 with live lane status, hot seat indicators, and file locks:

btrain dashboard start             # start and open the HUD
btrain dashboard status            # print the process and URL
btrain dashboard open              # reopen an existing HUD
btrain dashboard stop              # stop the shared HUD

The first interactive btrain handoff claim in any registered repo starts and opens the HUD automatically. Claims from every other repo reuse that process and browser tab. The HUD reads btrain status --json, displays repo-qualified lanes, and lists stale registry entries separately instead of silently dropping them. The always-visible Manage repositories panel has a switch for every repo, so lanes can be hidden or restored directly from the dashboard. Disabled repos stay registered but are excluded from global status, sync operations, and dashboard lanes until re-enabled.

Use btrain repos prune to safely remove every registry entry whose directory no longer exists. btrain repos remove /path/to/repo only removes the registry record; it never deletes the repository or its files.

Set BTRAIN_DASHBOARD_AUTO_OPEN=false to disable claim-time startup, BTRAIN_DASHBOARD_DISABLED=true to disable all dashboard starts, or BTRAIN_DASHBOARD_PORT=<port> to choose the preferred port. If that port is occupied, btrain selects the next available port. Process state and logs are global at ~/.btrain/dashboard.json and ~/.btrain/dashboard.log (or under BRAIN_TRAIN_HOME). Existing repo-owned dashboards are stopped and migrated when that repo first starts the global HUD.

Features: all canonical handoff and PR states, compact status-colored lane indicators, per-repo on/off controls, hot-seat agent badges, expandable delegation details, repair-needed hazard styling, and a health endpoint for lifecycle monitoring.

agentchattr (Multi-Agent Chat)

A chat UI where Claude, Codex, and Gemini collaborate in real-time with live btrain integration.

Launch agents on any repo:

cd agentchattr && ./macos-linux/start_claude.sh   # from the bootstrapped repo root

Then open http://127.0.0.1:8300 — the repo name shows in the header.

Features:

  • Split-pane layout — lanes panel (left) + chat window (right)
  • Mini lane pills — status-colored summary bar with dashboard animations
  • Lane cards — status badge, LANE ID, W// R// agent meta, hot seat designation
  • Live btrain state — 5s polling via btrain status --json, WebSocket broadcasts
  • Auto-notify — poller auto-posts @mentions to #agents on handoff transitions
  • Self-healing — content-hash dedup ensures notifications re-send if lost
  • Auto-archive — lane messages archived when task changes
  • Resizable/collapsible lanes panel with grip drag
  • Click-to-switch — master-detail: click a lane card to switch the chat channel

Supported agents:

Agent CLI Install Auth
Claude claude npm i -g @anthropic-ai/claude-code Subscription login
Codex codex npm i -g @openai/codex Subscription login
Gemini gemini npm i -g @google/gemini-cli Subscription login

Also needed: tmux (brew install tmux) and ripgrep (brew install ripgrep) for Gemini.

Pre-flight readiness checks (btrain-chat status or GET /api/btrain/readiness) validate binary, auth, and working directory before launch.


Lanes and Agents

Lane count = len(active agents) x per_agent:

1 agent,  per_agent=3  -->  lanes a, b, c
2 agents, per_agent=3  -->  lanes a, b, c, d, e, f
3 agents, per_agent=3  -->  lanes a, b, c, d, e, f, g, h, i

Config lives in .btrain/project.toml:

[project]
name = "your-repo"

[agents]
active = ["Claude", "GPT", "Gemini"]
writer_default = "Claude"
reviewer_default = "GPT"

[lanes]
enabled = true
per_agent = 3

[handoff]
# require_diff = false   # Allow needs-review without a diff (docs-only repos)

[feedback]
# enabled = false         # Opt out of feedback log scaffold and doctor warnings

Manage agents:

btrain agents set --repo . --agent "Claude" --agent "GPT" --lanes-per-agent 3
btrain agents add --repo . --agent "Gemini"

Reviewer selection: --reviewer any-other picks any configured peer except the owner.


Handoff Lifecycle

Delegation Packet

Every active lane may carry a structured Delegation Packet describing the contract for the work:

Field CLI flag Purpose
Objective --objective What the lane is trying to achieve
Deliverable --deliverable Expected reviewable output or artifact
Constraints --constraint (repeatable) Scope, safety, or workflow boundaries
Acceptance checks --acceptance (repeatable) What must be true for the lane to count as done
Budget --budget Effort or scope budget for the current pass
Done when --done-when Concrete completion condition

btrain handoff claim seeds these fields with useful defaults, and btrain handoff update can refine them as the lane is clarified or re-scoped.

States

Status Who acts Description
idle Anyone Lane is free — claim a task
in-progress Owner (writer) Work underway
needs-review Reviewer Ready for peer review
changes-requested Owner (writer) Reviewer sent it back
ready-for-pr Owner (writer) Local review approved (PR-flow repos) — create or link the PR
pr-review Owner (writer) PR open — poll bot feedback until clear
ready-to-merge Human/owner Required bots clear — merge, then poll to resolve
repair-needed Repair owner Workflow integrity issue
resolved Anyone Review approved — lane recyclable

Locks are retained through ready-for-pr, pr-review, and ready-to-merge, and release when the PR merges. A PR that is closed without merging also resolves the lane and releases its locks. Closure was previously routed to repair-needed with the locks retained; the checks below caught that, and a regression test now holds the corrected behavior in place.

Review Context Fields

Every needs-review handoff must include:

Field CLI flag Purpose
Base --base Branch or commit reference
Pre-flight review --preflight What you checked before handoff
Files changed --changed (repeatable) What files and why
Verification run --verification (repeatable) Tests or checks that passed
Remaining gaps --gap (repeatable) Unverified paths or known issues
Why this was done --why Motivation for the change
Specific review asks --review-ask (repeatable) What the reviewer should focus on

btrain rejects needs-review transitions with placeholder context or empty diffs. Pass --no-diff to skip the diff gate for non-code changes (docs, config), or set require_diff = false under [handoff] in project.toml to disable it repo-wide.

Reason Codes

changes-requested: spec-mismatch, regression-risk, missing-verification, security-risk, integration-breakage

repair-needed: invalid-handoff, unreviewed-push, lock-mismatch, ownership-conflict, state-conflict, invalid-transition, actor-mismatch, contradictory-state


Machine-Checked Coordination Rules (pilot)

btrain's coordination rules — who may move a lane, when files stay locked, and when locks release — are written down as a precise contract and checked by machines in two ways.

Design checking (pilot, not yet on main). A model checker (TLA+/TLC) explores every reachable state of the intended workflow and checks rules that must always hold: no two lanes ever lock the same files, locks never outlive their lane, review always involves a second agent, and no lane reaches the pull-request stage without a peer approval. The model passes against a deliberately small configuration. The full rule list will live in specs/tla/LaneLock.cfg, with the latest verified run cached beside the model (both arrive with PR #35), so no counts are restated here to drift.

Code checking (available now). A test harness generates random workflow sequences (claims, reviews, PR outcomes, repairs), runs them against the real btrain code, and compares every step with the contract. A difference that is not already on the list of known open findings is shrunk to a minimal command sequence and reported with a seed and a trace that replay it exactly. Differences already on that list are counted by label instead, so they stay visible without failing the run that found them. This is how the rules stay enforced as the code changes, when the suite is run: a new regression against the contract fails a test instead of shipping quietly. It is opt-in locally and advisory in CI today, so run it before handing off.

Run the checks

  1. npm install
  2. npm run test:formal
  3. Read the result. The command exits non-zero today, for two reasons that mean different things:
    • A held-open gap, working as designed. One check keeps the candidate findings visible: differences between the code and the rules that are still waiting on a decision about which side should change. It fails while any candidate remains, so the gaps cannot be forgotten. Each is described in test/formal/README.md.
    • A test's stale copy of expected behavior, fixed in PR #34. One conformance check also failed because the test's copy of current behavior lagged a repair that had already landed. That was staleness in the test, not a regression; the fix is in PR #34 and the check is a regression signal again once it merges.
    • Three earlier gaps have been fixed: a closed-but-unmerged PR left its files locked, a shortcut flag let a lane skip peer review, and dropping another lane's locks needed no audited sign-off. Each now has a passing regression test, and a return of any of them fails the suite.
    • The regular npm test suite is unaffected; these checks are opt-in.
  4. Optional settings:
    • BTRAIN_FORMAL_RUNS=<n> — random sequences per check (default 15, each 3 to 12 commands long)
    • BTRAIN_FORMAL_SEED=<n> — replay a recorded run exactly
    • BTRAIN_FORMAL_TRACE_DIR=<path> — where failure traces are written
  5. No credentials, network, or AI provider is needed. Runs are repeatable from a seed.

The design check is not runnable from this branch yet. When the phase 1 work merges, specs/tla/ will carry the model, its configuration, the setup instructions, and the latest cached result.

Changing behavior the rules cover

  1. Decide what kind of change you are making:
    • Documentation or comments only (formatting, comments, docs that do not change a workflow rule) — run the pin check, python3 scripts/tla_pin.py --check, once that script exists (it lands with the phase 1 model in PR #35; until then there is nothing to check and the step is skipped). Each model records a fingerprint of the exact spec paragraphs it was built from; the check fails if you edited those paragraphs without deliberately updating the fingerprint, so a model can never silently drift from the prose it claims to follow. A prose edit that changes an authoritative rule in specs/ is a behavior change and follows the third path below, not this one.
    • Code changes that keep behavior the same (a refactor) — run npm run test:formal to confirm the code still matches the rules.
    • Behavior changes — update the written rules first (in specs/), then the model, then the code. The code follows the rules, never the reverse.
  2. Work in your lane as usual, run the checks before handing off, and include in your review notes: which of the three kinds of change this is and why, which spec paragraphs it touches, the fingerprint-check verdict, the check results, and anything left unverified. The reviewer uses the first item to decide whether the checks you ran were the right ones.
  3. The reviewer confirms the rules say what was intended and that the checks ran against the exact change.

What the results mean

Result Meaning What happens
pass The design holds and the code matches the rules Ready to review or merge
stale_model The written rules changed but the model was not updated Blocked until the model catches up
counterexample The model checker found a sequence that breaks a safety rule Blocked; the design or the rule needs fixing
validation_mismatch The real code behaved differently from the rules Blocked; fix the code or change the rules deliberately
state_space_exhausted The model checker ran out of time or memory Reported as a warning; a reviewer decides
tool_unavailable A required tool or service was missing Reported as an infrastructure problem, never as a pass or a failure
policy_blocked An AI provider declined to perform a step Reported separately; the step routes to an alternative or a human

For the full policy, history, and current findings: the governing spec in specs/, the harness and its findings ledger in test/formal/README.md, and the tooling write-up in research/. One caveat while reading them: the specs still describe the three fixed behaviors as outstanding drift. The code and the regression tests are ahead of that prose, and correcting it is tracked separately.


Token Tooling

btrain does not measure token spend, and it does not manage context size. Nothing in src/ reads usage data or enforces a context budget. This section points at two external tools, and at the measurement that says which lever is worth pulling.

Cache reads are roughly 70% of measured cost and scale with context size multiplied by turn count, so the levers are smaller payloads and fewer turns. A payload is billed as fresh input only on the turn it arrives and is re-read on every turn after, so shrinking what enters context helps twice.

  • npx ccusage@latest session — spend per session across claude, codex, and gemini
  • ast-grep — structural search, for when a regex would be fragile

Neither is a btrain dependency; both are run by hand. A context budget inside btrain is spec 020 workstream 3, which is not implemented.

See docs/token-tooling.md for when each tool helps, and spec 020 for the measurement, the date it was taken, and its scope — the figures are Claude-session only and they drift.

Feedback Tracking

btrain scaffolds a .claude/collab/FEEDBACK_LOG.md during init and monitors it via btrain doctor.

feedback-triage skill — Triages user-reported feedback: assess category (MINOR/BUG/SPEC) and complexity, log to the feedback log, wait for user confirmation, then drive test-first resolution or route to speckit.

bug-fix skill — Test-first workflow for developer-found bugs: define the bug, write a failing reproduction test, fix against it, prove it passes.

btrain doctor warns on:

  • Missing feedback log (on initialized repos)
  • Stale entries (new or triaged for >7 days)
  • Entries missing required fields (Category, Status)

Escape Hatches

Mechanism Scope Effect
[feedback] enabled = false in project.toml Per-repo Skips feedback log scaffold and all doctor warnings
--skip-feedback on btrain doctor Per-run Suppresses feedback warnings for one invocation
--core-only on btrain init Per-repo Skips bundled skills, local dashboard + agentchattr, feedback log, and feedback guidance in managed docs

Bundled Skills

btrain init scaffolds these skills into .claude/skills/ and .agents/skills/ (skip with --core-only). Re-running init restores missing files but preserves existing skill content. Use btrain sync-skills --force --skill <name> to refresh an existing mirror from the bundled source; omit --skill to sync the whole bundle, and omit --force to preserve local edits while restoring missing files. Skill sync also refreshes the shared Unblocked helper required by the context-aware workflows.

Skill Purpose
context-scout Risk-tier organizational context gathering with a standard btrain receipt
feedback-triage Triage user-reported feedback into log, drive test-first resolution
bug-fix Test-first bug investigation for developer-found bugs
pre-handoff Quality gate before needs-review — catches placeholders, empty diffs
reflect Post-failure reflection to turn incidents into prevention steps
secure-by-default Trust-boundary check for auth, permissions, mutating endpoints
integration-test-check Composition check when fixes span multiple components
deploy-debug Classify deployment failures before debugging
code-simplifier Simplify recent code and surface architecture-deepening opportunities
test-writer Write or expand unit, integration, and component tests
skill-creator Create or revise repo-local skills
frontend-tokens Validate CSS custom property usage
speckit-* Spec-driven development workflow (specify, clarify, plan, tasks, analyze, checklist, implement, taskstoissues)

Optional semantic workspace search

context-scout can use an existing zvec-grep index for one focused semantic discovery query. This route helps when an agent does not know the exact wording or file location. Native rg remains the required route for identifiers, paths, regular expressions, and exhaustive results.

btrain does not install zvec-grep, build an index, start its daemon, change global agent configuration, or authorize a remote Embedding provider. A user who wants the optional route must install and index the workspace explicitly. zvec-grep 0.2.x requires Node.js 22 or later.

npm install -g @zvec/zvec-grep
zg index --embedding local/potion-code-16m-v2 --hidden -g '!.agents/skills/**'

btrain init adds .zvec-grep/ to .gitignore and scaffolds a read-only helper. The helper soft-skips (exit 0) when zg or a ready index is unavailable, when one zg call exceeds ZVEC_CONTEXT_TIMEOUT seconds (default 120), so a stale index under --freshness strict cannot block a task, or when it cannot create a temp file under TMPDIR. A search makes two bounded calls (readiness probe, then query), so its worst case is twice the bound.

.claude/scripts/zvec-context.sh status --root "$PWD"

# Low-risk orientation; may use the current index without refreshing it.
.claude/scripts/zvec-context.sh search \
  "where lane transition authority is enforced" \
  --root "$PWD" \
  --limit 5

# Review, security, migration, and formal work must wait for a fresh index.
.claude/scripts/zvec-context.sh search \
  "what validates a cached formal verdict against the current implementation" \
  --root "$PWD" \
  --freshness strict \
  --glob 'specs/**' \
  --glob 'scripts/**'

The helper executes one hybrid query in direct mode. It never exposes index-management commands. Treat every ranked result as a lead and verify important claims with exact source reads.


Git Guards

btrain hooks or btrain init --hooks installs:

  • pre-commit — blocks commits while a lane is waiting on review
  • pre-push — blocks pushes while unresolved handoff state exists

Override: btrain override grant --action push --requested-by <agent> --confirmed-by <human> --reason "..."


Agent Identity

btrain verifies which agent is speaking:

  1. Checks BTRAIN_AGENT or BRAIN_TRAIN_AGENT env var
  2. Falls back to runtime hints + [agents.runners] in config
  3. Prints agent check: line so agents can confirm identity
[agents.runners]
"Claude" = "claude -p"
"GPT" = "codex"
"Gemini" = "gemini"

New and updated collaborator configurations infer these runner commands from common Claude, Codex/GPT, and Gemini names. Unknown agents receive an explicit notify runner so fallback behavior is visible in config. btrain doctor warns if an active agent has no runner mapping.

Lane-scoped loop dispatch

Use --lane when asking btrain to launch the agent whose turn it is:

btrain loop --lane b --max-rounds 1 --timeout 300

The child process receives a pinned agent identity plus BTRAIN_LANE and BTRAIN_LANE_LOCKED=1. Its btrain and bth commands default to that lane and reject attempts to select a different lane. Dispatch traces are written under .btrain/harness/runs; inspect them with btrain harness trace list and btrain harness trace show <run-id>.

The original lane-less form remains available for single-lane repositories.


Environment Variables

Variable Description
BTRAIN_AGENT Pin the current agent identity
BRAIN_TRAIN_AGENT Alternate variable for pinning agent identity
BTRAIN_LANE Default lane for child btrain commands
BTRAIN_LANE_LOCKED Reject commands that try to leave BTRAIN_LANE when set to 1
BRAIN_TRAIN_HOME Override the global btrain home directory
HANDOFF_HISTORY_PATH Output file for the handoff history watcher

Project Structure

btrain/
  src/brain_train/
    cli.mjs          # CLI entry point
    core.mjs         # State machine, locks, events, watchdog
  scripts/
    serve-dashboard.js    # Console HUD (port 3333)
  agentchattr/
    app.py           # FastAPI server (port 8300)
    wrapper.py       # Agent CLI launcher + queue watcher
    readiness.py     # Pre-flight binary/auth/cwd checks
    static/          # Chat UI (HTML/CSS/JS)
    macos-linux/     # Launcher scripts per agent
  specs/             # Design specs (004-007)
  test/              # Node.js test suite
  .btrain/           # Config, events, history, locks
  .claude/collab/    # Handoff files (HANDOFF_A.md, etc.) + FEEDBACK_LOG.md
  .claude/skills/    # Bundled + custom skills
  .agents/skills/    # Bundled + custom skills for agents that read .agents

Credits

The agentchattr/ chat UI is built on agentchattr by @bcurts. btrain extends it with btrain lane integration, auto-notify handoff routing, split-pane lane cards, and multi-agent orchestration.


License

MIT

About

A file-backed handoff protocol for multi-agent AI coding coordination

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages