# Dictation — developer handoff **Superhairpieces branch portal. 14 September 2026.** Companion to `../HANDOFF.md`. Build that first — this fills the same fields it gates on. The stylist speaks while examining the client; the fields fill; the suggestion follows. Nothing new happens after the fields are filled — the panel already recalculates whenever its watched fields change, so dictation is simply another way of answering questions. ``` push-to-talk -> chips appear as you speak -> Add -> fields -> suggestion ``` Detection is **continuous**. Chips appear while the stylist is still talking, so they can see they are being heard rather than finding out at the end. **Add** is the commit: it is the only thing that writes to the form. Discard throws the whole recording away. --- ## The one rule this feature lives or dies by **The model proposes. The stylist confirms. Nothing is written by voice alone.** The engine refuses to guess on an incomplete form — that is R-INCOMPLETE, and it is deliberate: half of men's consultations are missing a core field, and a confident suggestion built on a half-filled form is worse than none. Dictation that quietly fills `remaining_thickness: average` from a mumbled half-sentence walks straight around that protection and produces a confident suggestion from something nobody said. It would be the most damaging bug in the product, and it would look exactly like the feature working. So every proposed answer must show **the verbatim phrase it came from**: > Thickness of remaining hair — **fine** > from *"what's left is fine"* That is what makes confirmation take two seconds instead of a field-by-field re-check. **A proposal without a verbatim phrase must be dropped, not shown.** Corollary: dictation **adds**, it never overwrites *silently* — and that last word matters. The first build of this dropped any field that was already answered. Testing caught what that feels like: the stylist says "Norwood four" over an existing `6`, and **nothing happens**. They cannot tell whether they were misheard, ignored, or not recorded at all. So a contradiction is surfaced as a **change**, defaulted to rejected: > Baldness level — **4** > from *"Norwood four"* > Already answered as **6** — accept only if you meant to change it The invariant worth holding: **absence of a proposal always means "not heard".** Never let it also mean "heard and discarded". --- ## The trigger A 62px rounded square, white on the form's own hairline border, with a mic glyph and the **AI badge riding the top-right corner** — so the control reads as "the AI one" before the mic is even parsed. Same indigo as every other AI surface, per `../HANDOFF.md` §6.7. It is icon-only, so it carries `aria-label="Dictate — fill answers by voice"` and the same as a `title`. An unlabelled mic button is guessable; it should not have to be guessed by a screen reader. Anchored **bottom-right**, 22px in, clear of the save bar — centred, it sat over the questions being answered and over the save bar's status text. Full width on screens under 760px. ## How a detection looks A detected answer is a **chip** — uppercase, outlined, with an ✕ to drop it: > `NORWOOD 6 ✕` `SWEATY SCALP ✕` `STRAIGHT ✕` `MEDIUM ON TOP ✕` > `FINE UNDERNEATH ✕` `WANTS NATURAL ✕` `AT OUR SALON ✕` `TAPE / GLUE ✕` Chips, not rows, because eight of them scan in about a second and a stack of eight rows does not. Each field says how to phrase itself (`SHP_DICT_CHIP` in `extract.js`) — "6" alone is meaningless and "Baldness level: 6" is too long to scan a row of. The provenance requirement survives the compression: **the transcript sits directly above the chips with every matched phrase highlighted.** The stylist checks the chips against their own words in one glance instead of expanding eight rows. Hovering a chip gives the full field name and phrase. An inferred chip ("he swims a lot" → sweaty) is **amber, not indigo**, and says so on hover. Removing a chip offers **Undo**. A mis-tapped ✕ on a busy salon floor should not mean re-dictating. ### Not every chip deserves the same attention Eight chips that all look equally certain means checking eight, and checking eight is slow enough to cancel out the reason for dictating in the first place. So the extractor distinguishes **what was said** from **what was interpreted**: | said | interpreted | |---|---| | "Norwood six" → `6` | "he swims a lot" → sweaty scalp | | "straight hair" → `Straight` | "wants something that lasts" → `durable` | | "about four inches" → `medium` | "six inches" → `medium` *(sits on the medium/long edge)* | Interpreted ones are **amber, sorted first, and counted in the header** — *"1 needs a look — the other 3 matched your words exactly"*. Verification goes from eight items to one. The boundary case is worth keeping: 6″ is the edge between a stock `medium` and a custom order. Any spoken measurement within 0.75″ of a bucket edge is flagged, because that is a judgement the machine should not make silently. This is the same lesson as `../HANDOFF.md` §6.3 — when everything looks equally important, nothing is. ### A contradiction is not a chip A chip says *this is going in*. A contradiction is a **question**, and it has to be asked in words: > You said "**Norwood four**", but Baldness level is already **6**. > `Keep 6` `Change to 4` Neither option is preselected, and it does not count toward "Fill N answers" until answered. Accepting turns it into a chip like any other. --- ## Before any of this ships: consent The consultation opens with a blocking privacy consent that **names the exact channels** client information passes through. This feature adds two that are not named today: 1. **Client voice audio leaves the building** to whichever transcription vendor you choose. 2. **Audio and transcript are retained** on the consultation record. Voice recording of a client is not covered by "we record your contact details, measurements and photos". It needs its own line in the consent, and a retention period. There is already the right pattern for it — the optional photo consent, which is separate from the blocking one and can be declined without stopping the appointment. Model this the same way: **dictation off by default, enabled per consultation once consent is given, and the stylist can still type.** Practical consequence: the dictation button must be **disabled and explained** when voice consent has not been given, not hidden. A stylist who cannot find the button will assume it is broken. --- ## Two endpoints ### 1. Transcribe ``` POST /api/consultation/transcribe multipart: audio blob + consultationId -> { transcript, durationMs, vendor, audioId } ``` Server-side, not the browser's Web Speech API. Three reasons, in order of weight: Web Speech streams audio to Google (Chrome) or Apple (Safari), which is a third party your consent does not name and you have no agreement with; it is Chrome/Safari only; and it is materially worse on accents and on product codes like "M111XT", which stylists say constantly. Budget roughly **a cent per consultation** and **1–3s** for a short utterance. ### 2. Extract ``` POST /api/consultation/extract-fields { transcript, alreadyAnswered: [field ids] } -> { proposals: [ { field, value, phrase, confidence } ], unrecognised: "..." } ``` Constrain the model's output to **`extraction-schema.json`** in this folder. That file is **generated from the form** — regenerate it whenever a question's options change, and never hand-edit it. Three constraints the handler must enforce server-side, not trust the model for: - **`value` must be a member of that field's enum.** Anything else is a rejected response, not a value to coerce. "Norwood 6ish" is not `6`. - **`phrase` must appear verbatim in the transcript.** Check it with `transcript.includes(phrase)`. If it does not, drop the proposal — a phrase the model paraphrased is a phrase the stylist cannot verify. - **Fields in `alreadyAnswered` come back marked, not dropped.** Return them with `conflict: { was }` so the UI can offer the change; the client defaults them to rejected. Dropping them here is what produced the silent-contradiction problem above. What must be enforced server-side is that a conflicting value cannot be written without an explicit accept — not that it is hidden. `confidence: 'low'` marks an inference rather than something stated — "he swims a lot" implying a sweaty scalp. Show those differently and do not pre-accept them. The prototype tints them and says *"Inferred, not stated"*. --- ## It is wired into the consultation prototype `../hairstylist-form-full-prototype.html` carries the frame, bottom-right, clear of the save bar. Two things about how it is wired are worth copying: **It writes the way a click writes.** `values[id]` for a single answer; `multi[id]` plus the joined string for a multi-select — then `render()`. The `render()` call is not optional: it is what refreshes the suggestion panel, the progress rail and the "N required left" counter. Setting state without it fills the boxes and leaves everything downstream stale. **The extractor is inlined, and generated.** The prototype is standalone by design — open the file, it works — so the extractor has to live inside it. An inlined copy is exactly the drift that hid three bugs in this project (`../HANDOFF.md` §8), so it is regenerated, never hand-edited: ``` python3 dictation/build-into-prototype.py # regenerate from extract.js python3 dictation/build-into-prototype.py --check # CI: fails if they diverged ``` Every class is prefixed `dct-` — the form already has its own `.chip`. Measured in the form itself: eight answers by voice take it to 10 of 17, and the panel still **refuses** — "Complete density, hairpiece type". A second utterance completes it and returns **M115 Lace, HD105 Lace**: no skin bases, because "scalp gets really sweaty" reached R-SWEAT through the voice path. ## What is in this folder | File | What it is | |---|---| | `extract.js` | **a stand-in**, see below | | `extraction-schema.json` | generated from the form; the model's allowed output | | `build-into-prototype.py` | regenerates the inlined copy, `--check` for CI | The working demo is **in the consultation prototype** — open `../hairstylist-form-full-prototype.html` and press the mic. There was a separate standalone demo here; it was removed once the frame landed in the form, because two demos of the same flow is the drift this project has already been bitten by three times. `git log` has it if you want the isolated version back. **`extract.js` is not what ships.** It is a keyword matcher that produces the same *shape* as the model would, so the prototype demonstrates the flow and the contract without a key or a network call. Do not extend it; replace it. What is worth keeping from it is the **contract** at the top of the file and the `alreadyAnswered` behaviour. There is no microphone in the prototype either — clicking an example replays it word by word into the frame, standing in for live transcription. The suggestion underneath **is** the real engine. --- ## What the prototype demonstrates, measured One spoken sentence: > "He's a Norwood six, scalp gets really sweaty, straight hair. Wants about four > inches on top. What's left is fine. He wants it to look as natural as possible > and he'll come to us for maintenance. Tape and glue." - Chips land **as the words arrive** — 1 at four words, 3 at eight, 4 at sixteen. - **8 answers** by the end, each with its phrase highlighted in the transcript. - Pressing **Add** writes all 8 and the save bar moves from "7 required left" to "4 required left" — and the panel **still refuses**: *"Complete density, hairpiece type."* The gate holds against voice exactly as against typing. This is the single most important thing to preserve. - A second utterance — "Toupee, medium density." — completes it, and the suggestion appears. One thing worth copying rather than rediscovering: because detection reruns on every word, the chip entrance has to be scoped to chips that are **actually new**. Animating all of them on each pass leaves them at opacity 0 most of the time — they flicker, and a screenshot catches an empty box. And the hard rule survives the voice path, which is the thing worth testing: | same client | shortlist | |---|---| | dictated "scalp gets really sweaty" | M154 Hybrid, M105 Lace — **no skin** | | same client, not sweaty | M101 Skin, M101V Skin, M105 Lace | R-SWEAT fired from spoken words. That is the check to automate. --- ## Definition of done - [ ] Voice consent exists, is separate from the blocking consent, and can be declined without blocking the appointment. - [ ] With consent not given, the button is **visibly disabled with a reason**. - [ ] Every proposal shows a phrase that appears verbatim in the transcript. - [ ] A value outside the field's enum is rejected server-side, not coerced. - [ ] A field the stylist typed is never overwritten by voice *without an explicit accept* — and a contradiction is shown, never silently dropped. - [ ] Dictating a partial consultation leaves the panel **refusing and naming what is missing** — it must not become answerable by voice alone. - [ ] A dictated sweaty scalp returns **zero skin bases** (R-SWEAT through the voice path). - [ ] Rejecting a proposal leaves that field untouched, not blank-filled. - [ ] `extraction-schema.json` regenerates from the form, and a changed option set fails loudly rather than silently producing invalid values. --- ## Measured on the prototype's own logic Running the extractor over the first example, then over a contradicting one: | check | result | |---|---| | answers proposed from one sentence | 8 | | every proposal's phrase appears verbatim in the transcript | yes | | contradictions surfaced rather than dropped | 3 of 3 | | contradictions defaulted to rejected | 3 of 3 | | genuinely new answers defaulted to accepted | 3 of 3 | ## The part that needs no verification The transcript itself is kept, verbatim, in **`hairstylist_internal_note`** — opt-out, checked by default. This is the safest and probably the most valuable half of the feature. Free text has no enum to violate, nothing to silently mis-fill, and nothing to check: it is the stylist's own words, unaltered. That field is filled on roughly half of consultations today and read by nothing, so the reasoning behind a choice is being lost on every other client. If the structured extraction turns out not to survive real speech, **this part still earns its place** — and it is a fraction of the work. Worth shipping first for that reason. (The prototype did not render `hairstylist_internal_note` at all until this landed, though it exists on the live form. Added as "Your note".) ## Known risks - **Homophones on product codes.** "M115" / "M150" / "M155" are a transcription coin-flip. Do not let dictation set `recommendedProducts` — the stylist picks systems from the panel and the search, as now. - **Two people talking.** The client speaks during the examination too. Whatever the client says about their own preferences is useful, but attributing it is not solved here; treat the transcript as one voice in v1. - **Numbers.** Base measurements dictated as "seven by nine" are easy to mis-hear and hard for the stylist to spot. Consider keeping measurements typed. - **Silence is not "no".** An unmentioned field must stay unanswered. Never map absence onto a default — that is the R-INCOMPLETE bypass again, wearing a different hat.