LAB RECORD / J-009 / 17-20 AUGUST 2026
Evopien now has a voice baseline worth freezing.
This record covers the engineering work after the latest website commit on 17 August 2026 at 16:22 CEST. The production voice path now reaches a representative clean physical response of about 2.15 seconds from user end-of-speech to first meaningful audio. The fastest observed turn in the retained evidence was about 1.51 seconds.
Those numbers must not be collapsed into a guaranteed 1.5-second claim. Healthy physical turns still vary, and the frozen master checkpoint describes an ordinary clean range of roughly 2.0 to 2.7 seconds. The honest public milestone is therefore about two seconds, with 1.51 seconds recorded as a best observed turn rather than the normal ceiling.
Current result
What is frozen, active and still open.
Voice owns acoustic input and output. Conversational Presence is a separate active subsystem that will decide social timing without changing the frozen voice baseline.
Local bilingual full duplex, endpointing, meaningful first audio, streaming speech and semantic interruption form the production baseline.
Representative clean physical result after user end-of-speech. The fastest observed retained turn was about 1.51 s; this is not a guarantee.
Parakeet TDT 0.6B v3 Q4_K runs as a persistent six-thread CPU worker; same-WAV median was about 0.297 s.
Selective-INT8 Piper uses Amy Medium for English and Claude High for Spanish, about 166 MB RSS and roughly 0.20 s synthesis for typical short opening pieces.
Qwen3.5-4B UD-Q3_K_XL runs through llama.cpp with reasoning off, 4096 voice context, Q8 K/V and mandatory prompt-cache warm-up.
Turn state, thinking pauses, floor handoff and acknowledgement are being evaluated separately from Voice V1.
No production mhm, mm-hm or uh-huh backchannel is enabled yet.
Direction of arrival remains observe-only and no robust family-room floor owner is claimed.
Selective governed memory, the complete multi-user privacy proof, face or head actuation and physical movement remain future milestones.
Engineering sequence
What changed after J-008.
The work was not one model swap. It was a measured narrowing of the complete response path until the remaining latency no longer justified destabilising the subsystem.
- The response boundary was defined honestly
Timing is measured from physical user end-of-speech to the first acoustic answer that carries real meaning. Filler such as wait, let me think or a synthetic acknowledgement added only to stop the clock does not count.
- Parakeet became a persistent six-thread service
Q4_K CPU recognition remained inside the 8 GB memory envelope while same-WAV median time fell from about 0.764 s at two threads to about 0.297 s at six. GPU ASR remained rejected because its unified-memory risk was not justified.
- The endpoint stopped letting random noise own the clock
The normal pause remains about 0.60 s, but production capture updates its final speech time only for qualified speech after the turn opens. Environmental noise no longer extends a turn merely because it crosses a raw energy threshold.
- English and Spanish TTS candidates were tested, then narrowed
Scylla, Pocket and other alternatives exposed useful speed and quality trade-offs, but no mixed stack justified continuing model churn. Selective-INT8 Piper with Amy Medium and Claude High became the final bilingual Voice V1 baseline.
- Qwen was profiled beyond headline tokens per second
Prompt cache removed roughly 2.6 to 2.8 seconds from cold prompt evaluation. The resident multimodal projector added about 55 MB PSS and essentially no measured text latency. Batch size, K/V type, pinned clocks and context depth produced only small gains.
- Latency shortcuts were rejected when they damaged the product
Speculative rolling paths reached roughly 1.26 to 1.68 s but were fragile. Forced ultra-short first sentences could reach roughly 1.11 to 1.64 s but degraded answer quality and gamed the metric. Neither became production Voice V1.
- The final physical voice result became good enough
Representative recent live turns clustered around 1.95, 2.12 and 2.18 seconds, averaging about 2.08 seconds. The frozen master checkpoint records a representative median near 2.15 seconds and an ordinary clean range around 2.0 to 2.7 seconds.
- Voice V1 was tagged and frozen
The suite reached 233 passing tests plus two subtests. Bash syntax, Python compilation and whitespace validation passed. Further voice tuning now requires a concrete physical-prototype failure or a changed product requirement.
- Conversational Presence became the active branch
The next problem is not simply speaking earlier. It is deciding whether silence means finished, thinking or continuing, and whether a subtle acknowledgement can show attention without taking the floor.
Frozen production baseline
Voice V1 on the 8 GB Jetson.
These are the selected production choices. Earlier J-008 details remain historically correct for that checkpoint but no longer describe the active TTS or ASR lifecycle.
- Compute
- Jetson Orin Nano Super 8 GB · MAXN_SUPER · Ubuntu 24.04.4 · JetPack 7.2.1-b49 · CUDA 13.2
- LLM / VLM
- Qwen3.5-4B UD-Q3_K_XL · llama.cpp · GPU offload · 4096 voice context · Q8 K/V · reasoning off · multimodal projector resident
- ASR
- Parakeet TDT 0.6B v3 Q4_K · persistent native CPU worker · six threads · RIGHT XVF channel
- TTS
- Piper selective INT8 · Amy Medium English · Claude High Spanish · first piece near 48 characters plus slack
- Audio
- XVF3800 hardware AEC · LEFT speech gate · RIGHT ASR · Pebble V3 playback · full-duplex semantic interruption
- Endpoint
- Approximately 0.60 s normal pause · qualified-speech timing · no random-noise extension
- Default vision
- Fresh C920 frame on demand · resident multimodal projector · continuous RF-DETR/NvDCF off
- Preserved vision
- DeepStream 9.1 · RF-DETR Medium 576 FP16 · threshold 0.40 · NvDCF tracking · separate validated mode
Measured timing
The two-second milestone, without hiding its variance.
Component timings explain the architecture, but only physical end-to-end timing supports the user-facing response claim.
Best observed first meaningful audio in the retained 18-turn evidence; not the representative result.
About 1.95 s, 2.12 s and 2.18 s after the stronger direct-answer prompt and final voice path.
Clean serial physical end-of-speech to first meaningful audio; ordinary healthy turns were roughly 2.0 to 2.7 s.
Endpoint confirmation, WAV finalisation and ASR consume most of the time before Qwen can begin.
Representative warm first-sentence median from the retained Qwen profile, including first-token time.
Typical short opening piece in the frozen selective-INT8 production path.
Engineering decision
A lower benchmark number is no longer the highest-value task.
Q4_K_M decoded about 11.6% faster than Q3 in isolation, but the expected opening-audio saving was only tens of milliseconds. Larger batches saved about nine milliseconds. Pinned Jetson clocks added heat for about a quarter-percent decode gain. None changes the interaction as much as correct turn timing.
Sub-second meaningful speech would require safe overlap between live ASR, end-of-turn confidence, speculative Qwen and TTS. That remains a legitimate future architecture, but it is not justified while Voice V1 already meets the current prototype requirement and major robot systems remain unfinished.
Active subsystem
Presence owns social timing, not the answer.
The objective is for Evopien to feel genuinely there: listening, waiting when appropriate, acknowledging carefully and responding at the socially correct moment.
The user is still speaking or has clearly left the sentence open. Evopien remains silent.
Silence is not automatically a handoff. The user may still own the conversational floor.
A future cached mhm or subtle visual nod may signal attention without agreeing with the content or ending the turn.
The user has handed over the turn and frozen Voice V1 may generate the actual response.
A human mhm, yeah or okay while Evopien speaks should not always trigger barge-in.
Turn-detector checkpoint
Two small local detectors fit. Neither has earned floor ownership.
Smart Turn v3.2 and LiveKit v1-mini both deploy on the Jetson with acceptable inference cost. Smart Turn used about 176 MB peak RSS and later ran near 88 ms median. LiveKit used about 359 MB peak RSS and ran near 85 ms median. Compute feasibility passed for both.
On a deliberately difficult guided stress corpus at the 600 ms decision point, Smart Turn produced 21 false interruptions across 23 HOLD events; LiveKit produced 23 across 23. A threshold strict enough to remove false interruptions caused 39 unnecessary waits across 41 END events for either detector.
Evidence boundary
What this milestone does not claim.
1.51 s is the fastest retained observation. Representative clean operation is about 2.15 s and harder turns can take longer.
The architecture and benchmark path are defined, but backchannels and floor-control influence are not enabled in production.
A neutral acknowledgement must never be presented as agreement, factual verification or safety approval.
Presence cannot fabricate user turns, write long-term memory, bypass governance or issue raw motor commands.
Governance, selective memory, household privacy, face and gaze, movement, manipulation and real-world embodiment remain incomplete.

