Shared decision making, measured while it happens
ListenAide transcribes a clinical encounter in real time and re-scores the whole conversation against all five Observer OPTION⁵ items after every turn. It produces two deliberately distinct outputs — a Live Coaching Score during the visit and a Validated Observer Score after it — and cites the specific turns behind every item score.
End-to-end demonstration: live scoring interface, post-encounter review, exported report.
The measurement problem
Shared decision making is the standard of care. Measuring it hasn't kept up.
Shared decision making (SDM) — the process by which clinician and patient jointly reach a care decision informed by evidence and by what matters to the patient — is a stated priority of every major quality framework. It is also one of the least measurable behaviors in medicine.
The reference standard, Observer OPTION⁵, requires two trained raters to review a complete recorded encounter. Agreement between calibrated humans is moderate (ICC = 0.67), throughput is roughly one encounter per rater-hour, and every score arrives long after the conversation it describes has ended.
The consequence isn't merely inconvenience. It's the constraint that shapes what the field can ask — and, because the instrument returns a single terminal number, the sequential structure at the heart of SDM theory is collapsed and discarded at the moment of measurement.
ListenAide treats Observer OPTION⁵ not as an endpoint to be automated but as a computational representation of SDM behavior: a way of rendering a conversational process into a machine-readable signal that evolves across an encounter. The instrument is grounded in the "talk model" (Elwyn et al., 2012), which describes decision conversations as a progression:
What the bottleneck forecloses:
Training
A clinician who wants to improve receives, at best, a score and a general impression weeks later. Nothing identifies which moments succeeded or failed, or what a better version would have sounded like.
Quality improvement
SDM cannot sit on a quality dashboard beside measures that are extracted automatically. Health systems that value it have no instrument that scales to their volume.
Science
Questions about how deliberation unfolds and where it breaks down are unanswerable — not because they're uninteresting, but because no instrument returns process data.
The measure
Five validated items. One encounter-level score.
Observer OPTION⁵ (Elwyn et al., 2013) scores five SDM behaviors, each on a 0–4 effort scale, summed and rescaled to 0–100. ListenAide scores the encounter, not the speaker — crediting patient-initiated behaviors when the clinician responds supportively, exactly as the manual specifies. Watch how items come online as a conversation progresses:
composite = round(Σ items / 20 × 100)
How it works
Glanceable during the encounter. Auditable after it.
Mid-conversation, a clinician can't study a dashboard. The live interface is built so two things land in under a second: "How am I doing?" and "What should I do next?"
Browser-based speech capture
Both participants' speech is transcribed in the browser and displayed in a unified transcript, labeled simply Clinician and Patient. No special hardware — a laptop in the room is enough.
Cumulative scoring, every turn
After each turn, the backend re-analyzes the entire conversation so far and updates the Live Coaching Score. The best observed effort per item is retained, following the manual's guidance for scoring across encounters — early weaker attempts never drag the score down.
One nudge at a time
A large score ring and a single contextual prompt — "Consider: have you explained the pros and cons?" — surface the first unaddressed item, following the talk model's natural sequence. Nudges guide; they never alter scores.
Dual-endpoint architecture
Two scores, two validity claims — and never one number wearing two hats.
The Live Coaching Score and the Validated Observer Score are deliberately separate constructs. Formative in-encounter feedback need not carry the measurement properties of a validated instrument, and a validated instrument need not be available in real time. Holding both, and never conflating them, is a design position — the prototype enforces it in the interface, in the exported report, and in the data model.
Live Coaching Score
Validity claim: none. Formative process data.A provisional, continuously updated score of the encounter so far — built for in-the-moment guidance, not for measurement.
- Re-computed from the full transcript after every turn
- Best-per-item accumulation, mirroring the manual's guidance
- Drives the score ring, the talk-model phase badge, and the sequential nudge
Validated Observer Score
Mirrors the task the instrument was validated for.A separate, dedicated analysis of the complete transcript as a single holistic rating — the same task a trained rater performs from a full recording.
- Prompted with the complete Observer OPTION⁵ manual: every item, every anchor
- Computed independently of the Live Coaching Score — free to raise or lower any item
- Returns a per-item rationale and the specific turns cited as evidence
Post-encounter review
Every score comes with its evidence.
Automated scoring in clinical contexts that can't be interrogated will not be trusted, and shouldn't be. When the session ends, a three-view review panel opens the encounter up for audit — for the clinician's own reflection, supervisor review, or a training portfolio. The full encounter also exports as a formatted PDF report or a machine-readable JSON record for research and archival.
Audit the evidence, turn by turn
Every turn that contributed to any item's score is tagged. Click a turn to see exactly which items it supported and at what effort level; filter by item to isolate the evidence base behind any single score.
The measure scores the collaborative process — so patient-initiated SDM counts when the clinician responds supportively.
Rationale you can read, evidence you can check
Each item expands into its Validated Observer Score, a written rationale, the specific turns cited as evidence — rendered inline with speaker and text — and a comparison against the Live Coaching Score explaining any change.
See the encounter's temporal structure
The Live Coaching Score is plotted across every turn, with overlays for each item — showing when each SDM dimension first came online and how the encounter built. A gold marker plots the Validated Observer Score against the final live value, making any delta between the formative signal and the validated endpoint visually explicit.
The scientific opportunity
Because scoring is continuous, each encounter yields a trajectory — not a number.
A trained rater integrates a whole conversation into five numbers, and the sequence is lost. Turn-level re-scoring keeps it. Every existing application of Observer OPTION⁵ produces one score after the encounter ends; this produces a different data object entirely.
- Time to first observation — the turn at which each item first scores above zero.
- Item ordering — observed sequence against the talk model's expected progression.
- Accumulation slope — how quickly the composite builds, early versus late.
- Dwell time — the proportion of the encounter spent in each talk-model phase.
- Terminal delta — the distance between the final Live Coaching Score and the Validated Observer Score.
Illustrative only. Whether these trajectories cluster into stable, recoverable digital SDM phenotypes is an open empirical question and a proposed aim of the research below — not a finding. The curves at right are schematic.
Options and deliberation support named up front; the encounter plateaus high.
Little SDM behavior until the decision is imminent, then a rapid climb.
Strong information exchange, preferences never elicited or integrated.
Evidence base
Built on a validated instrument and published automation research.
Inter-rater reliability of Observer OPTION⁵ among trained human raters, with intra-rater reliability of r = 0.93. This is the benchmark automated scoring has to be measured against — not perfection.
Barr et al., 2015 · Patient Educ Couns
Correlation of LLM-generated OPTION⁵ scores with trained human raters — roughly 75–80% of the human–human agreement of r = 0.77. Anchored scored exemplars materially improved performance.
Elwyn, Selvaraj, Yen & Forcino, 2025
ListenAide preserves the instrument's architecture in both outputs: five items, five effort anchors each, one encounter-level composite rescaled to 0–100.
Elwyn, Grande & Barr, 2018 manual
What we claim — and what we don't
The psychometric properties of Observer OPTION⁵ were established for a post-encounter rating made by trained humans. They do not transfer automatically to real-time, automated scoring. The Live Coaching Score is formative process data and makes no measurement claim. The Validated Observer Score reproduces the rating task for which the instrument was validated — a single holistic judgment of a complete encounter. Whether it reproduces trained-rater ratings is an empirical question, and answering it against dual calibrated raters is Aim 1 below.
Known departures are explicit by design: automated rather than dual-human rating; no timestamp-based segmentation of encounters containing more than one decision; and the automation literature's consistent finding that the preference items (4 and 5) are the hardest to score — for machines and for humans alike.
Research agenda
The prototype exists so the questions can be asked. Three aims, with thresholds set in advance.
ListenAide is preliminary work for a proposed exploratory project in computational shared decision making. Each aim carries a success criterion specified before data collection — including the criteria that would count as failure.
Measurement validity
Score a corpus of transcribed primary care encounters twice: with the ListenAide holistic engine, and with two trained human raters calibrated to ICC > 0.60 and blind to the automated output. Prompt configurations vary systematically, each run in replicate to characterize scoring stability.
Digital SDM phenotypes
Extract trajectory features from turn-level score histories and test whether they cluster into recurring behavioral patterns. Stability is the crux, tested by split-half replication, bootstrap resampling, and leave-one-clinician-out analysis to confirm phenotypes describe encounters rather than individual habits.
Point-of-care feasibility
Human-factors evaluation with clinicians and advanced trainees in standardized-patient encounters, plus silent-mode deployment in live clinical encounters — the system scores but displays nothing, isolating technical feasibility from behavioral influence.
The work is exploratory and carries real risk. Automated agreement may fall short of human benchmarks, and phenotypes may prove unstable. Either would be a substantive result the field currently lacks: an instability finding would indicate that SDM trajectories are dominated by encounter-specific content rather than recurring behavioral patterns, redirecting effort from taxonomy toward encounter-level modeling.
Nothing in the approach is specific to a disease, specialty, or care setting. The unit of analysis is the decision conversation, which occurs in oncology, primary care, surgery, and palliative medicine alike — primary care encounters are the model system, not the target.
Architecture & privacy
A local-first system. Clinical conversations never have to leave the room.
- Local LLM analysis. Scoring runs on a locally served language model (via Ollama), so transcripts can be analyzed without any cloud dependency — a deliberate choice for clinical privacy.
- FastAPI backend, SQLite persistence. A lightweight server records sessions, turns, per-turn item scores, and the persisted Validated Observer Score — reopening a review never re-runs the analysis.
- Graceful degradation. If the model is unavailable, the backend returns safe fallback scores and the interface remains fully usable in a demonstration mode.
- Open, exportable data. Every session exports as a formatted PDF report for clinical and training use, and as complete JSON — turns, timestamps, the full Live Coaching Score history, and the Validated Observer Score object — for research pipelines.
Investigators
Measurement science and computational modeling, in the same room.
Computational modeling, LLM systems, and the ListenAide prototype architecture.
Shared decision making science; co-author of the original Observer OPTION⁵ instrument (Elwyn et al., 2013).
Judgment and decision making; analysis of deliberative process and choice behavior.
Scientific grounding
References
See the full encounter lifecycle in three minutes.
From the first spoken turn to the Live Coaching Score, the Validated Observer Score, and the evidence-cited export — the demonstration walks through everything on this page, live.