Why voice agents need ITN
Speech recognition can return spoken-form text, while downstream systems often expect written forms. A tool often needs a structured value instead:seven eight three two nine becomes an order ID, and twenty dollars becomes an
amount. A formatting mistake can become a failed lookup or an invalid tool
argument.
Learn where ITN fits in a voice-agent pipeline.
How Premove works
What it can write
These are supported candidate forms, not a promise that every sentence will
select that candidate. Context determines the final output. See
supported forms and boundaries.
Measured results
Premove ITN scored 398/400 (99.50%) on the voice-agent subset of its frozen synthetic stress benchmark. Its overall semantic entity accuracy was 89.70% across the 1,500-row suite. Mean warm latency was 56.49 ms on the measured Apple M4/MPS setup. The benchmark is not a production-traffic accuracy estimate; Premove was also slower than both comparison backends. See the full comparison, method, and limitations.Install and try it
normalize_structured() and provide a
NormalizationContext when a deterministic temporal value is required. The
same inference result powers readable text, resolved text, and span metadata.
The first use downloads about 1.6 GB of model files. The package pins the
matching release snapshot, so upgrading the package selects the new model on
the next initialization. Load one instance and
reuse it for later requests. The frozen-model release is validated on Apple
Silicon MPS and Linux x86-64 CPU; see deployment for the
complete platform boundary.