MEASURED, NOT CLAIMED
Every figure below comes from one stated machine. The comparison that matters is the lower two rows below: same measurement mode, so the difference between them is contention and nothing else.
Rows two and three are measured the same way, so the gap between them is the queue and nothing else: background conversation costs roughly 5×. Row one is the path a shipping game takes, where speech begins at the first finished sentence instead of at the end of the reply. No amount of tuning removes the 5× — throttle background conversation while the player is speaking.
The prompt cache is still filling. It happens once per session, and it is measured rather than excused.
Twenty sessions of sixteen turns on the reference machine. She is a reliable witness to what she was told — ask her to compute something and that is a different story.
State descriptors reach the line about six times in ten. With the system off: zero.
THREE THINGS THAT ARE DIFFERENT
FULLY OFFLINE
Model, recognition, synthesis and memory all run in the player’s process. Pull the network cable mid-conversation and nothing changes. No key to provision, no rate limit, no per-token bill, no vendor who can deprecate your characters.
MULTILINGUAL, WITH THE FINE PRINT
The three layers do not cover the same set of languages, so here is each one separately. One voice ships, and it is English — public domain, nothing to clear. Everywhere else the model and the recognition are already there and you supply the voice file: one asset, no code.
ALREADY SHIPPED IN A VR GAME
This is not a demo looking for a product. A shipped VR title runs on it, with players talking to characters on their own hardware — which is where the numbers on this page come from.
PERCEPTION, OFF AND ON
Same scene, same question, same character. The only difference is whether she can see the room she is standing in.
She invents a fireplace, a window seat and a cat. None of them exist.
She speaks only about what is in the room.
WHAT IT DOES NOT DO
Three limits worth knowing before the purchase rather than after it.
DESKTOP ONLY
It needs a desktop GPU and desktop memory. There is no phone build and no cloud fallback to hide behind — that absence is the point of the product.
ONE INFERENCE QUEUE
Every character shares one queue on one GPU. A tavern of six people all answering at once is not what this does; a scene with one or two speaking characters is. The 12.7 s figure above is that contention, measured.
PROBABILISTIC OUTPUT
It will not deliver an authored line word for word. If a beat has to land exactly, script that line and let the model handle everything around it. Threshold events, traits and save / load are the deterministic half.
GET IT
Paid once. No account, no key, no metering — there is no server to meter you against.
BUY THE PLUGIN THE STORE PAGE OPENS WHEN UNITY’S REVIEW COMPLETES. THE LINK IS ALREADY FINAL.