Agent / LLM Engineer · Applied AI · AI Memory & Execution · Multimodal Systems · Full-Stack AI
I build reliable AI systems that connect agent orchestration, long-term memory, permission-aware execution, multimodal context, evaluation, and production engineering.
Resume · LinkedIn · Hugging Face · Nexora · Owl Intelligence
- Owl Intelligence — Co-founder, Full-Stack AI Engineer & Product Lead
Architecting Klik, an AI memory-and-execution wearable platform across React/TypeScript, Python/FastAPI, PostgreSQL/Redis/Elasticsearch/pgvector, APISIX, Docker, KMP clients, and C/C++ embedded software. - Nexora — Founder, Agentic AI Systems & R&D
Building agentic workspaces, governed multi-agent collaboration, permission-scoped knowledge and execution, long-horizon workflows, and safe context handoffs.
| Project | Focus |
|---|---|
| OhMyHerdr | Native Rust agent-workspace runtime for Codex, Pi, and Claude Code profiles, with persistent instructions, isolated workspaces, and PTY-based shared terminals. |
| Ornith 256K LocalMoE | Reproducible Apple Silicon sparse-MoE serving study with verified 260,013-token input and agent-task performance comparisons. |
| CLI-Bench | Benchmark for AI agents learning and orchestrating command-line tools through deterministic stateful environments. |
| Soundscape | Live-API music and content discovery platform with geolocated rankings, listening telemetry, accessibility, browser tests, and atomic deployment. |
| Cheat Engine CLI | JSON-first Go CLI for authorized process inspection and guarded memory workflows on macOS, Windows, and remote ceserver targets. |
- User-scoped typed facts with source-session provenance, confidence, and confirmation tracking
- Asynchronous embeddings and weighted Elasticsearch BM25 + pgvector reciprocal-rank fusion
- Version-aware correction, scoped deletion, derived-index reconstruction, and LoCoMo/LongMemEval/ConvoMem evaluation
- On 1,986 LoCoMo QA pairs, hybrid retrieval improved R@5 from 0.356 to 0.611 (+25.5pp)
- Versioned approvals, durable action ledgers, credential and integration boundaries, and auditable external-write receipts
- Seatbelt/Docker sandboxing, failure recovery, and reusable best-practice/anti-pattern context
- Toolathlon: 56/63 tasks passed (88.9%); sandbox reuse reduced cold start from 16s to 150ms
- Qwen3-ASR Flash/Volcengine ASR, local Whisper, Pyannote 3.1 diarization, and ECAPA voice embeddings
- Name/voiceprint/speaker reconciliation through user-scoped typed entities and a 19-case reconciler
- Nine-factor session-boundary detection without interrupting continuous recording
- KLIK Temporal, Entity-Aware, Privacy-Constrained Memory
- KLIK-Bench: Memory-Grounded Multi-Tool Orchestration
- CLI-Bench: Command-Line Tool Orchestration
- Small Active Expert, Big Context Window
My earlier work spans LLM, agent, and AI-wearable research at Minerva Capital, strategic investment at Zhipu AI, and TMT equity research. Training in finance and econometrics informs how I approach statistics, evaluation, product economics, and measurable outcomes.
Agent / LLM Engineering, Applied AI, AI memory and retrieval, agentic execution, multimodal AI, AI infrastructure, full-stack AI, and selective remote, contract, and AI-evaluation opportunities.
Contact: wilson_xu@panor.tech · Bilingual resume



