LAB RECORD / J-008 / 14-17 AUGUST 2026
The first integrated Evopien head now hears, speaks, sees on demand and measures where speech comes from.
This record covers the verified development window after the previous website checkpoint: 14 August at 15:16 CEST through the latest verified commit on 17 August at 01:45 CEST. In that period, the prototype moved from a working local voice loop to an integrated bilingual head with fresh camera turns, a preserved continuous perception mode, and the first measured source-attention experiments.
The result is substantial but bounded. Evopien can converse and answer current visual questions locally. It cannot yet reliably decide which of several nearby people owns the conversational floor, and the 8 GB unified-memory budget still creates serious latency and swap pressure under the heaviest configurations.
Current result
What the head can do now - and where the claim stops.
The current default is a stationary conversational head. Heavy continuous perception is preserved as a separate mode for future movement, navigation and manipulation.
Qwen3.5-4B, Parakeet and Kokoro run locally on the Jetson.
XVF3800 hardware AEC allows listening and interruption during playback.
Normal turns and substantive barge redirects can switch languages in both directions.
A visual request captures a current C920 JPEG, attaches it to the same model turn and releases the camera.
DeepStream 9.1, RF-DETR Medium and NvDCF run together, but remain off in the default stationary-head launcher.
XVF angle data is recorded as supporting evidence and is not allowed to reject speech.
A lightweight speaker-continuity signal must be measured before audio source fusion is enabled.
Persistent governed memory, the full multi-user privacy proof and physical actuation remain future milestones.
Chronology
The 14-17 August engineering sequence.
The order matters because each result changed the architecture of the next experiment.
- 14 August - current-frame visual turns entered Evopien Core
Provider-neutral model messages gained text and image parts. The voice path learned to detect visual intent, capture one 1280×720 MJPEG frame from the C920 with FFmpeg, keep the JPEG in memory, send it through llama.cpp's multimodal API and fail explicitly when capture was unavailable. This was on-demand vision, not continuous sight.
- 15 August - the acoustic topology was corrected
Speaker output was routed Jetson USB → XVF3800 → 3.5 mm AUX → Pebble V3. This gives the XVF3800 the exact far-end reference its hardware AEC expects. LEFT/channel 0 became the clean near-end gate; RIGHT/channel 1 remained the intelligible ASR signal.
- 15 August - interruption became conversational instead of phrase-bound
Playback now pauses before semantic confirmation when the clean human gate qualifies a candidate. Adaptive endpointing classifies control, complete, continuing and uncertain speech while microphone capture continues during interim ASR. Valid first candidates are preserved instead of being overwritten by a second capture.
- 15 August - the stack became bilingual and faster to hear
Parakeet TDT 0.6B v3 Q4_K replaced the larger Q8 checkpoint; Kokoro ONNX FP16 replaced the slower PyTorch/Piper path for the active bilingual runtime; English
am_michaeland Spanishem_alexbecame the selected voices. Substantive interruptions were validated from Spanish to English and English to Spanish, while short controls intentionally preserved the current language. - 15 August - environmental impacts stopped acting like commands
Silero VAD was inserted after the hardware-clean energy gate. Keys, mouse movement, doors and many cutlery impacts produced near-zero speech probabilities and no longer paused playback. A control latch made an accepted Wait survive later blank or noisy endpoint probes. Stop remains usable but less reliable because very short ASR tokens can still be corrupted.
- 16 August - the 8 GB memory envelope was measured, not assumed
Headless operation, disabled Snap services, 2 GiB ZRAM, a 4 GiB NVMe swap safety net, Q8 K/V cache and an 8192-token Qwen context were benchmarked. A populated 6009-token prompt fit without process swap in the isolated Qwen test, while the complete voice stack peaked near 5.9 GB.
- 16 August - continuous NVIDIA perception became real
DeepStream 9.1 was integrated with C920 hardware MJPEG decoding. YOLO26n became the control, RF-DETR Nano was exported and parsed, and RF-DETR Medium 576 FP16 at threshold 0.40 with detector interval 5 and NvDCF tracking was selected from synchronized measurements.
- 16 August - the two vision modes were separated deliberately
The continuous service can publish
evopien.scene.v1state and a shared live JPEG without a second camera owner. The default head instead keeps RF-DETR and NvDCF off, captures a fresh frame only when the spoken turn requires vision, and leaves the proven continuous mode available for later embodiment. - 17 August - the default Qwen quant changed, but the optimization proof remains open
The head launcher moved from Q4_K_M at 1536 context to Qwen3.5-4B UD-Q3_K_XL at 8192 context. The integrated head started and remained functional, but this checkpoint does not claim that the new quant has already passed the full 30-minute and endurance memory-acceptance targets.
- 17 August - direction data was tested and prevented from becoming a false authority
XVF3800 DoA polling was integrated into the real conversation loop as telemetry. Controlled SOUTH/EAST testing proved that direction is useful but can remain stably wrong for several consecutive samples. Directional speech rejection was therefore not enabled.
Current default head
The selected stationary-head runtime.
These are the active architectural choices at the final checkpoint, not a list of every experiment attempted along the way.
- Compute
- Jetson Orin Nano Super 8 GB · MAXN_SUPER · JetPack 7.2.1-b49 · L4T R39.2.1 · CUDA 13.2
- LLM / VLM
- Qwen3.5-4B UD-Q3_K_XL · 8192 context · Q8 K/V · GPU offload · non-thinking · multimodal projector available
- ASR
- Parakeet TDT 0.6B v3 Q4_K · CPU · transient process; GPU mode rejected after CUDA allocation failures
- TTS
- Kokoro ONNX FP16 · persistent HTTP service · am_michael / em_alex
- Audio
- XVF3800 hardware AEC · LEFT near-end gate · RIGHT Parakeet channel · SYS_DELAY 12 · REF_GAIN 8 · MIC_GAIN 90
- Turn admission
- Silero speech qualification · Parakeet transcript admission · 600 ms normal endpoint
- Default vision
- Fresh C920 frame on demand · continuous RF-DETR/NvDCF off
- Preserved vision
- DeepStream 9.1 · RF-DETR Medium 576 FP16 · threshold 0.40 · interval 5 · NvDCF tracking
- Direction
- XVF3800 DoA observe-only telemetry; no direction-based speech rejection
Measured evidence
The benchmark numbers that changed the design.
Isolated component numbers and full-stack numbers are kept separate because combining them would create a false memory budget.
Mean about -65.7 dB, maximum about -30.8 dB; the far-end phrase was removed from the clean gate.
Maximum RMS about 0.7214; the clean near-end signal strongly exceeded the 0.040 gate.
2027.9 MB average RAM, 15.0% average GPU, 6.62 W average, 57.7°C average TJ; tracker held about 30 FPS.
6502.3 MB average RAM, 7048 MB peak, 197 MB maximum swap, 34.4% average GPU, 66.2°C maximum TJ.
Qwen processed a real C920 frame; 977 prompt + 250 completion tokens took 20.39 seconds.
A 395,212-byte 1280×720 JPEG was read while DeepStream retained sole ownership of /dev/video0.
About 1.18-1.29 seconds versus roughly 0.06-0.12 seconds from the continuous shared camera.
Includes nine new source-angle tests; no compile errors and clean whitespace validation.
Primary systems blocker
The head works, but the heaviest baseline was swapping hard.
The full on-demand multimodal head profiler reached about 6504 MB average RAM and 7376 MB peak on a 7485 MB usable system. Minimum MemAvailable fell to about 16.2 MB, swap peaked around 2.70-2.73 GB, and peak swap-out reached about 577,512 KB/s. This is not harmless reserve use; it is memory thrashing and a direct contributor to conversational latency.
CPU Parakeet probes ranged from roughly 1.8 to 7.8 seconds, sometimes above 10 seconds, and multi-probe turns could exceed 15 seconds. GPU Parakeet was rejected after unified-memory allocations of roughly 396 MiB failed. Long Kokoro replies could also require 10-18 seconds of synthesis work. Qwen text itself was often much faster when memory pressure was not dominant.
Decisive negative result
Direction of arrival cannot identify the speaker by itself.
Catalin speaking from the SOUTH bench position produced very stable readings around 158° with Silero probabilities near 0.99-1.00. A second speaker physically at EAST was initially tracked around 103°, 81°, 73°, 69° and 68°.
During the same EAST-speaker stretch, the XVF then reported a sustained SOUTH-like cluster: 170°, 184°, 191° and repeated 192° values before returning toward EAST. The problem was therefore not one random bad sample. Room reflection, multipath or beamformer state can create several consecutive stable but physically incorrect angles.
Known limits
What this milestone does not prove.
No production speaker-embedding model or similarity threshold has been selected.
Dog, baby, TV and other people can still create contextual ambiguity even when non-speech impacts are filtered.
Parakeet remains transient and is a major latency contributor.
Stale llama-server processes caused one startup OOM; service-death detection and child cleanup need hardening.
The local 4B model can hallucinate facts and visual details; retrieval and verification remain future work.
Temporary session context is not the selective, inspectable, privacy-separated memory system defined by the Foundation.
The ESP32 exists, but no motor protocol, safety watchdog, joint controller or physical action milestone is integrated.
Demonstration media
The demo video is intentionally still pending.
A public recording will be added after the current head scenario is stable enough to show normal conversation, interruption, English/Spanish switching, fresh visual questions and explicit known limits without editing failures out of the engineering record.

