Skip to content
View tombudd's full-sized avatar

Block or report tombudd

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
tombudd/README.md

Tom Budd

AI Safety & Governance Engineer · Research-Grounded Systems Builder

Website Research LinkedIn Get Involved


I build small, reproducible evaluation artifacts for researchers and engineers reviewing agentic AI systems.

My public work focuses on:

  • agent authority boundaries
  • tool-use safety
  • human oversight
  • auditable decision evidence
  • deterministic synthetic evaluations

Public repositories contain synthetic, clean-room, or educational artifacts. They do not expose private production systems, logs, prompts, schemas, or customer data.


Start Here

Evidence-Gated Evaluation of Agentic AI Systems — METHODS PREPRINT / RUNNABLE / TESTED

A bounded framework for testing agent capability without expanding authority. The public release includes the manuscript, JSON Schema, 18 synthetic cases, deterministic gate logic, regression tests, frozen results, and independent-review materials.

Start with:

frozen claim + authority envelope + evidence ledger -> PASS / HOLD / FAIL / INVALID_RUN

A passing result supports only the predeclared claim within the tested environment. It grants no authority to deploy or act.

Evaluation benchmark portfolio

AI Governance Benchmarks — RUNNABLE / TESTED

A clean-room benchmark suite using synthetic cases, deterministic scoring, generated reports, and regression tests.

synthetic case -> scorer -> report -> tests

Start with:

What the portfolio demonstrates

  • Runnable synthetic evaluation cases
  • Deterministic scoring and reproducible reports
  • Tests that detect boundary failures and unsupported public claims
  • Explicit separation between capability evidence and operational authority

Additional Public Work

  • Agent Action Audit Template — RUNNABLE / TESTED — schema-backed synthetic action receipts, blocked-action examples, human-review metadata, and validation tests.
  • Human-AI Governance Lab — RUNNABLE / TESTED — toy workflow gates for risk classification, human approval, reports, and synthetic audit receipts.
  • Active Inference Primer — RESEARCH_NOTES / UTILITIES — minimal educational free-energy utilities with synthetic numerical tests and explicit limitations.
  • Eudaimonic Alignment — RESEARCH_NOTES — public research notes on human flourishing, agency, and alignment/governance questions.
  • Quantum AI Experiments — SANDBOX — simulator-first quantum/AI-adjacent experiments with explicit claim boundaries.

Contact

I am open to serious collaborators in AI evaluation, agent safety, governance engineering, red-teaming, and applied research.

tombudd.com · tom@tombudd.com · Get involved

Pinned Loading

  1. agent-action-audit-template agent-action-audit-template Public

    Clean-room public template for synthetic agent action audit records and human review checkpoints.

    Python 1

  2. human-ai-governance-lab human-ai-governance-lab Public

    Clean-room project exploring reviewable, testable, accountable human-AI collaboration workflows.

    Python 1

  3. tombudd tombudd Public

    1

  4. eudaimonic-alignment eudaimonic-alignment Public

    Eudaimonic alignment — AI wellbeing as an alignment strategy, drawing from Aristotelian philosophy. CC BY 4.0.

    1

  5. active-inference-primer active-inference-primer Public

    Active inference for anticipatory cognition in autonomous AI agents. CC BY 4.0.

    Python 1