Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]
-
Updated
Mar 5, 2025 - Python
Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]
Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.
Reproducing agentic KV-cache policy claims on real traces. 68k requests from 393 Claude Code sessions. LRU is harder to beat than the papers suggest.
Reproducible llama.cpp CPU inference profiling and a deterministic LLM serving simulator with continuous batching, KV cache, prefix caching, and workload-driven latency analysis.
Models Take Notes at Prefill: KV Cache Can Be Editable and Composable (arXiv:2606.17107) — paper, code, results, and interactive companion.
Miniature LLM serving runtime with continuous batching, chunked prefill, KV-cache management, preemption, prefix caching, and real Llama-family execution.
GenPark AI Agent Skill - Radix tree prefix caching simulator for prompt templates and agent system instructions, maximizing KV-cache hit rates and cutting TTFT latency.
GenPark AI Agent Skill - Task-aware prompt token pruning using information density scoring and semantic stop-phrase elimination to compress long context prompts up to 60%.
GenPark AI Agent Skill - Task-aware prompt token pruning using information density scoring and semantic stop-phrase elimination to compress long context prompts up to 60%.
GenPark AI Agent Skill - Mixture-of-Agents (MoA) multi-layer aggregator synthesizing diverse candidate proposals from heterogeneous sub-agents into high-consensus outputs.
GenPark AI Agent Skill - Multi-tier model cascading router dynamically selecting optimal models (e.g. Small vs Medium vs Large) based on task complexity, cost budgets, and SLA constraints.
GenPark AI Agent Skill - Radix tree prefix caching simulator for prompt templates and agent system instructions, maximizing KV-cache hit rates and cutting TTFT latency.
GenPark AI Agent Skill - Speculative decoding simulation engine coordinating small draft model speculative token generation and large target model parallel rejection sampling verification.
GenPark AI Agent Skill - Speculative decoding simulation engine coordinating small draft model speculative token generation and large target model parallel rejection sampling verification.
GenPark AI Agent Skill - Mixture-of-Agents (MoA) multi-layer aggregator synthesizing diverse candidate proposals from heterogeneous sub-agents into high-consensus outputs.
GenPark AI Agent Skill - Multi-tier model cascading router dynamically selecting optimal models (e.g. Small vs Medium vs Large) based on task complexity, cost budgets, and SLA constraints.
Adaptive Disaggregated Inference on a Role-free Fleet!
Unified execution runtime for LLM and ML programs.
Cache-aware router for OpenAI-compatible LLM servers, in Go. Per-worker radix trees route each request to the worker holding its KV prefix. Validated on 4x A100 + vLLM and Apple Silicon + llama.cpp.
C++ inference runtime for llama.cpp that shares a single document KV-cache prefill across multiple analytical branches via snapshot fan-out. Eliminates redundant GPU compute and dramatically reduces TTFT in DAG-based multi-agent pipelines.
To associate your repository with the prefix-caching topic, visit your repo's landing page and select "manage topics."