# peripheral — Neuron's I/O organ (own-core, local, consent-gated) The interface made physical. Two afferent senses in, one efferent voice out — all reached the way the agentic surface reaches any tool. ``` MIC (hear) afferent device -> capture -> descriptor -> ingest -> geometry CAMERA (see) afferent device -> capture -> descriptor -> ingest -> scene-geometry SPEAKER(speak) efferent render WAV -> PLAY ALOUD out the speaker ``` Closes the conversational loop: **hear (mic) -> understand (engram) -> speak (speaker)**. ## Rails - **Own-core.** macOS-native only: AVFoundation (camera/mic), CoreAudio voice- processing (AEC), afplay (speaker), ImageIO/CoreGraphics (frames), hand-rolled DSP (WAV, LPC, formant synthesis). No cloud, no heavy deps. - **Local-only.** Raw streams are written to `out/` and never egress. `.gitignore` keeps captured media out of git. - **Consent-gated (two locks).** A Neuron-level grant (`grant`/`revoke`) *and* the OS TCC permission. Sensitive senses (camera/mic) fail closed without both. - **Disclosed.** Every device touch prints a `[peripheral]` line on stderr. ## Build ``` swiftc -O -o bin/periph src/periph.swift \ -framework AVFoundation -framework CoreMedia -framework Foundation \ -framework CoreGraphics -framework ImageIO -framework CoreImage ``` ## Commands ``` periph grant|revoke # Neuron-level consent periph status periph speak # SPEAK ALOUD (efferent) periph tone [hz] [sec] # own-core WAV synth periph listen # MIC capture (afferent), 16k mono periph see # CAMERA one frame (afferent) periph feat-audio | feat-image # capture -> compact descriptor periph ingest-audio|ingest-image # descriptor -> engram node (geometry) periph voiceprint # extract F0 + formants F1-F5 periph imitate # speak back in that voice (LPC resynthesis) periph hear-imitate # MIC -> signature -> imitate -> SPEAK ALOUD periph converse [--authority F] [--barge-at S[:backchannel|:bargein]] [--resume] [--live-mic] ``` ## The afferent metabolism A capture is never shipped raw. It becomes a **compact descriptor** — the afferent twin of the music instrument-signature: - audio -> `[seconds, sr, ch, rms, peak, zcr, centroid, F0]` (~2400-6000x smaller) - image -> `[w, h, meanRGB, brightness, 3x3 luminance grid]` (~400000x smaller) - voice -> `[F0, F1..F5, bandwidths]` (11 numbers) That descriptor is what the ingest organ (engram `POST /api/nodes`) turns into an embedded node = geometry. ## Voice by imitation `voiceprint`/`imitate` are own-core LPC (autocorrelation + Levinson-Durbin, order 16 @ 16 kHz), formant extraction from the LPC spectral envelope, and source-filter resynthesis (glottal impulse train at F0 through the all-pole formant filter). A voice is grabbed by ear as ~a dozen numbers and spoken back — **no training, no stolen voice.** Measured fidelity on real speech: resynthesized formants match the source within 2-3%. The full phoneme->formant path for *novel* sentences is the speech faculty's seam (`elp` audio surface profile); this engine provides the formant synthesis primitive it renders through. ## Interruptibility (native turn-taking) `converse` plays the utterance as an ordered, salience-tagged **meaning-plan** while the mic listens (full-duplex, AEC on so it never barges in on its own voice): - **barge-in**: user speech -> pause on the spot (sample-accurate), not "finish the buffer." - **yield-or-hold**: a decision grounded in the current segment's salience + progress + the interrupter's authority — YIELD (stop) or HOLD ("hang on, let me finish"). - **backchannel** ("mm-hm"): brief/low -> keep going, resume seamlessly. - **resumable**: on yield the remaining plan persists (`.resume.json`); `--resume` picks the thread back up ("as I was saying"). Live full-duplex uses `--live-mic` (OS AEC). Injected `--barge-at` drives the decision loop deterministically for testing. ``` ```