Forward deployed AI engineer · lead IC
I work with clients to put document AI and model inference into production, and I build the backend around it.
Remote from UTC-4. English and Spanish. Open to senior and lead IC roles.
th3nolo.com · Engineering notes · Open-source log · CV · LinkedIn
Most of my client work is private, so the results below link to case studies and write-ups. The public projects further down are code you can clone and run.
| System | Result | Source |
|---|---|---|
| Financial document AI, Arabic and English audited statements | 34% → 94% on an internal extraction and calculation evaluation. One 22-page filing in 97 s for USD 0.30 to 0.42 | case study |
| semgate: permission gate for coding agents, TypeSafe Jev judge | 0 of 47 harmful-allow base cases on a 355-case set (95% upper bound 6.2%). 83.0% of benign actions auto-allowed (95% CI 78.4% to 87.5%) | evals |
| LangGraph RAG chat, JEV-first routing | 0.981 vs 0.957 accuracy on 322 labeled messages | note |
| Small-model fine-tune (LoRA), training-data repair | Valid plans 1/81 → 75/81, strict pairwise F1 0.01 → 0.88 on 81 development cases | note |
| GLM-5.2-504B on 8× RTX PRO 6000 (vLLM, SM120) | 240K-token context after a sparse-attention layer fix | note |
| Project | What it does | |
|---|---|---|
| semgate | Permission gate for coding agents. Deterministic rules decide the certain cases, a typed Jev judge handles the rest, and every doubt goes to the human. Hooks into Claude Code, Codex, OpenCode, Gemini CLI, Antigravity, Copilot CLI, and more. 2,958 tests and TLA+ models. Host hook behavior measured with hookconf. | Python |
| quake-reunite | Search API and map for people and aid centers after the June 2026 La Guaira earthquake. Mistral OCR and Gemma 4 on Cerebras read WhatsApp lists, hospital-list photos, and PDFs, then merge duplicate people across sources. Live | Python |
| openrouter-mcp | Stateless MCP 2026-07-28 server and CLI for OpenRouter. 13 tools, 1 resource, HTTP and stdio. | TypeScript |
| sqlbench-harness | Benchmark for LLM-generated SQL on BIRD, KaggleDBQA, Defog SQL-Eval, and Spider 2.0. Treats benchmark text as untrusted prompt input. Reports accuracy, errors, tokens, and cost. | Python |
| verifiable-exchange-demo | Limit-order engine. Anyone can replay its history from signed orders, Merkle proofs, and on-chain anchors. Live demo | Rust |
| dep-age-gate | Refuses dependency versions younger than 72 hours. Lockfile audit, pre-commit hook, and GitHub Action. | Python |
- My PR #85 to the OpenClaw installer was merged (commit bfc0bd9). It rewrote the Windows install checks, and three of those functions still run in the live
install.ps1. - In hiero-sdk-js, I showed that the published
@hashgraph/sdkstill pulled in protobufjs 8.0.0 (GHSA-xq3m-2v4x-88gg, a remote code execution bug) after the upstream fix. The issue was closed as completed. - In openai/codex#34801, I traced broken image thumbnails to the desktop app fetching signed URLs without the Bearer token.
- I also reproduced 19 bugs in Codex, Claude Code, and OpenCode. The open-source log has the evidence for each one.
- How routing a LangGraph RAG chat with JEV got me 98% accuracy · 2026-09-27
- What I learned fine-tuning a small model: mistakes you can avoid · 2026-09-16
- Building a strict stateless MCP 2026-07-28 server for OpenRouter · 2026-08-31
- Building a Verifiable Exchange: Signed Orders, Replayable Execution, and On-Chain Anchors · 2026-08-27
- Serving GLM-5.2-504B on RTX PRO 6000: the vLLM sparse attention fix on SM120 · 2026-07-04
Python · TypeScript · Rust · Go · Solidity. FastAPI, NestJS, PostgreSQL, vLLM, Docker, GitHub Actions



