Working prototype · Computational shared decision making

Shared decision making, measured while it happens

ListenAide transcribes a clinical encounter in real time and re-scores the whole conversation against all five Observer OPTION⁵ items after every turn. It produces two deliberately distinct outputs — a Live Coaching Score during the visit and a Validated Observer Score after it — and cites the specific turns behind every item score.

Live Coaching Score · formative Validated Observer Score · post-encounter Auditable by construction

End-to-end demonstration: live scoring interface, post-encounter review, exported report.

The measurement problem

Shared decision making is the standard of care. Measuring it hasn't kept up.

Shared decision making (SDM) — the process by which clinician and patient jointly reach a care decision informed by evidence and by what matters to the patient — is a stated priority of every major quality framework. It is also one of the least measurable behaviors in medicine.

The reference standard, Observer OPTION⁵, requires two trained raters to review a complete recorded encounter. Agreement between calibrated humans is moderate (ICC = 0.67), throughput is roughly one encounter per rater-hour, and every score arrives long after the conversation it describes has ended.

The consequence isn't merely inconvenience. It's the constraint that shapes what the field can ask — and, because the instrument returns a single terminal number, the sequential structure at the heart of SDM theory is collapsed and discarded at the moment of measurement.

ListenAide treats Observer OPTION⁵ not as an endpoint to be automated but as a computational representation of SDM behavior: a way of rendering a conversational process into a machine-readable signal that evolves across an encounter. The instrument is grounded in the "talk model" (Elwyn et al., 2012), which describes decision conversations as a progression:

Phase 01 · Justify Team talk Explain that a decision needs deliberation; form a partnership to support the work.
Phase 02 · Inform Option talk A two-way exchange of high-quality information and opinions about the available options.
Phase 03 · Elicit & integrate Decision talk Listen for the patient's preferences and goals; integrate them as decisions are made or deferred.

What the bottleneck forecloses:

No feedback loop

Training

A clinician who wants to improve receives, at best, a score and a general impression weeks later. Nothing identifies which moments succeeded or failed, or what a better version would have sounded like.

Doesn't scale

Quality improvement

SDM cannot sit on a quality dashboard beside measures that are extracted automatically. Health systems that value it have no instrument that scales to their volume.

No process data

Science

Questions about how deliberation unfolds and where it breaks down are unanswerable — not because they're uninteresting, but because no instrument returns process data.

The measure

Five validated items. One encounter-level score.

Observer OPTION⁵ (Elwyn et al., 2013) scores five SDM behaviors, each on a 0–4 effort scale, summed and rescaled to 0–100. ListenAide scores the encounter, not the speaker — crediting patient-initiated behaviors when the clinician responds supportively, exactly as the manual specifies. Watch how items come online as a conversation progresses:

OPTION⁵ · live item detail Live Coaching Score 0/100
1
Alternate options existSignal that a decision with options is needed
2
Support deliberationReassure the patient they'll be supported in deliberating
3
Describe options & pros/consInform and check understanding across reasonable options
4
Elicit preferencesDraw out the patient's preferences about the options
5
Integrate preferencesFold elicited preferences into the decision being made
Effort scale: 0 none · 1 minimal · 2 moderate · 3 skilled · 4 exemplary composite = round(Σ items / 20 × 100)

How it works

Glanceable during the encounter. Auditable after it.

Mid-conversation, a clinician can't study a dashboard. The live interface is built so two things land in under a second: "How am I doing?" and "What should I do next?"

01 · Listen

Browser-based speech capture

Both participants' speech is transcribed in the browser and displayed in a unified transcript, labeled simply Clinician and Patient. No special hardware — a laptop in the room is enough.

02 · Analyze

Cumulative scoring, every turn

After each turn, the backend re-analyzes the entire conversation so far and updates the Live Coaching Score. The best observed effort per item is retained, following the manual's guidance for scoring across encounters — early weaker attempts never drag the score down.

03 · Guide

One nudge at a time

A large score ring and a single contextual prompt — "Consider: have you explained the pros and cons?" — surface the first unaddressed item, following the talk model's natural sequence. Nudges guide; they never alter scores.

Dual-endpoint architecture

Two scores, two validity claims — and never one number wearing two hats.

The Live Coaching Score and the Validated Observer Score are deliberately separate constructs. Formative in-encounter feedback need not carry the measurement properties of a validated instrument, and a validated instrument need not be available in real time. Holding both, and never conflating them, is a design position — the prototype enforces it in the interface, in the exported report, and in the data model.

During · Point of care

Live Coaching Score

Validity claim: none. Formative process data.

A provisional, continuously updated score of the encounter so far — built for in-the-moment guidance, not for measurement.

  • Re-computed from the full transcript after every turn
  • Best-per-item accumulation, mirroring the manual's guidance
  • Drives the score ring, the talk-model phase badge, and the sequential nudge
After · Validated rating task

Validated Observer Score

Mirrors the task the instrument was validated for.

A separate, dedicated analysis of the complete transcript as a single holistic rating — the same task a trained rater performs from a full recording.

  • Prompted with the complete Observer OPTION⁵ manual: every item, every anchor
  • Computed independently of the Live Coaching Score — free to raise or lower any item
  • Returns a per-item rationale and the specific turns cited as evidence

Post-encounter review

Every score comes with its evidence.

Automated scoring in clinical contexts that can't be interrogated will not be trusted, and shouldn't be. When the session ends, a three-view review panel opens the encounter up for audit — for the clinician's own reflection, supervisor review, or a training portfolio. The full encounter also exports as a formatted PDF report or a machine-readable JSON record for research and archival.

Audit the evidence, turn by turn

Every turn that contributed to any item's score is tagged. Click a turn to see exactly which items it supported and at what effort level; filter by item to isolate the evidence base behind any single score.

The measure scores the collaborative process — so patient-initiated SDM counts when the clinician responds supportively.

Rationale you can read, evidence you can check

Each item expands into its Validated Observer Score, a written rationale, the specific turns cited as evidence — rendered inline with speaker and text — and a comparison against the Live Coaching Score explaining any change.

See the encounter's temporal structure

The Live Coaching Score is plotted across every turn, with overlays for each item — showing when each SDM dimension first came online and how the encounter built. A gold marker plots the Validated Observer Score against the final live value, making any delta between the formative signal and the validated endpoint visually explicit.

The scientific opportunity

Because scoring is continuous, each encounter yields a trajectory — not a number.

A trained rater integrates a whole conversation into five numbers, and the sequence is lost. Turn-level re-scoring keeps it. Every existing application of Observer OPTION⁵ produces one score after the encounter ends; this produces a different data object entirely.

  • Time to first observation — the turn at which each item first scores above zero.
  • Item ordering — observed sequence against the talk model's expected progression.
  • Accumulation slope — how quickly the composite builds, early versus late.
  • Dwell time — the proportion of the encounter spent in each talk-model phase.
  • Terminal delta — the distance between the final Live Coaching Score and the Validated Observer Score.

Illustrative only. Whether these trajectories cluster into stable, recoverable digital SDM phenotypes is an open empirical question and a proposed aim of the research below — not a finding. The curves at right are schematic.

100 50 0 turns → score
Early establishing
Options and deliberation support named up front; the encounter plateaus high.
Late deliberating
Little SDM behavior until the decision is imminent, then a rapid climb.
Option-heavy
Strong information exchange, preferences never elicited or integrated.

Evidence base

Built on a validated instrument and published automation research.

ICC 0.67

Inter-rater reliability of Observer OPTION⁵ among trained human raters, with intra-rater reliability of r = 0.93. This is the benchmark automated scoring has to be measured against — not perfection.

Barr et al., 2015 · Patient Educ Couns

r ≈ 0.6

Correlation of LLM-generated OPTION⁵ scores with trained human raters — roughly 75–80% of the human–human agreement of r = 0.77. Anchored scored exemplars materially improved performance.

Elwyn, Selvaraj, Yen & Forcino, 2025

5 × 0–4

ListenAide preserves the instrument's architecture in both outputs: five items, five effort anchors each, one encounter-level composite rescaled to 0–100.

Elwyn, Grande & Barr, 2018 manual

What we claim — and what we don't

The psychometric properties of Observer OPTION⁵ were established for a post-encounter rating made by trained humans. They do not transfer automatically to real-time, automated scoring. The Live Coaching Score is formative process data and makes no measurement claim. The Validated Observer Score reproduces the rating task for which the instrument was validated — a single holistic judgment of a complete encounter. Whether it reproduces trained-rater ratings is an empirical question, and answering it against dual calibrated raters is Aim 1 below.

Known departures are explicit by design: automated rather than dual-human rating; no timestamp-based segmentation of encounters containing more than one decision; and the automation literature's consistent finding that the preference items (4 and 5) are the hardest to score — for machines and for humans alike.

Research agenda

The prototype exists so the questions can be asked. Three aims, with thresholds set in advance.

ListenAide is preliminary work for a proposed exploratory project in computational shared decision making. Each aim carries a success criterion specified before data collection — including the criteria that would count as failure.

Aim 01

Measurement validity

Score a corpus of transcribed primary care encounters twice: with the ListenAide holistic engine, and with two trained human raters calibrated to ICC > 0.60 and blind to the automated output. Prompt configurations vary systematically, each run in replicate to characterize scoring stability.

Prespecified successComposite ICC ≥ 0.60 against consensus human ratings · ≥ 80% accuracy classifying high- vs. low-SDM encounters
Aim 02

Digital SDM phenotypes

Extract trajectory features from turn-level score histories and test whether they cluster into recurring behavioral patterns. Stability is the crux, tested by split-half replication, bootstrap resampling, and leave-one-clinician-out analysis to confirm phenotypes describe encounters rather than individual habits.

Prespecified success≥ 3 phenotypes recovered with split-half adjusted Rand index ≥ 0.50 and interpretable correspondence to talk-model constructs
Aim 03

Point-of-care feasibility

Human-factors evaluation with clinicians and advanced trainees in standardized-patient encounters, plus silent-mode deployment in live clinical encounters — the system scores but displays nothing, isolating technical feasibility from behavioral influence.

Prespecified successSystem Usability Scale ≥ 70 · encounter duration increase ≤ 1 minute · bounded per-turn latency and dropped-turn rate

The work is exploratory and carries real risk. Automated agreement may fall short of human benchmarks, and phenotypes may prove unstable. Either would be a substantive result the field currently lacks: an instability finding would indicate that SDM trajectories are dominated by encounter-specific content rather than recurring behavioral patterns, redirecting effort from taxonomy toward encounter-level modeling.

Nothing in the approach is specific to a disease, specialty, or care setting. The unit of analysis is the decision conversation, which occurs in oncology, primary care, surgery, and palliative medicine alike — primary care encounters are the model system, not the target.

Architecture & privacy

A local-first system. Clinical conversations never have to leave the room.

  • Local LLM analysis. Scoring runs on a locally served language model (via Ollama), so transcripts can be analyzed without any cloud dependency — a deliberate choice for clinical privacy.
  • FastAPI backend, SQLite persistence. A lightweight server records sessions, turns, per-turn item scores, and the persisted Validated Observer Score — reopening a review never re-runs the analysis.
  • Graceful degradation. If the model is unavailable, the backend returns safe fallback scores and the interface remains fully usable in a demonstration mode.
  • Open, exportable data. Every session exports as a formatted PDF report for clinical and training use, and as complete JSON — turns, timestamps, the full Live Coaching Score history, and the Validated Observer Score object — for research pipelines.
API surface
GET/statusSystem readiness and model availability
POST/session/{user}/{client}Create an encounter session
POST/analyze/{user}/{client}Score a turn against the cumulative transcript — updates the Live Coaching Score
POST/review/{user}/{client}/{session}Generate the Validated Observer Score and holistic review
GET/export/{user}/{client}/{session}Complete machine-readable session export
GET/history/{user}/{client}List recent sessions

Investigators

Measurement science and computational modeling, in the same room.

Michael McCormick, PhD Contact PI

Computational modeling, LLM systems, and the ListenAide prototype architecture.

Maka Tsulukidze, MD, PhD, MPH Co-Investigator

Shared decision making science; co-author of the original Observer OPTION⁵ instrument (Elwyn et al., 2013).

Todd McElroy, PhD Co-Investigator

Judgment and decision making; analysis of deliberative process and choice behavior.

Scientific grounding

References

Barr, P. J., O'Malley, A. J., Tsulukidze, M., Gionfriddo, M. R., Montori, V., & Elwyn, G. (2015). The psychometric properties of Observer OPTION⁵, an observer measure of shared decision making. Patient Education and Counseling, 98(8), 970–976.
Elwyn, G., Frosch, D., Thomson, R., Joseph-Williams, N., Lloyd, A., Kinnersley, P., et al. (2012). Shared decision making: A model for clinical practice. Journal of General Internal Medicine, 27(10), 1361–1367.
Elwyn, G., Tsulukidze, M., Edwards, A., Légaré, F., & Newcombe, R. (2013). Using a "talk" model of shared decision making to propose an observation-based measure: Observer OPTION⁵ Item. Patient Education and Counseling, 93(2), 265–271.
Elwyn, G., Grande, S. W., & Barr, P. J. (2018). Observer OPTION⁵ Manual. The Dartmouth Institute for Health Policy and Clinical Practice.
Elwyn, G., Selvaraj, S. P., Yen, R. W., & Forcino, R. C. (2025). Automating the Observer OPTION-5 measure of shared decision making: Assessing validity by comparing large language models to human ratings. Patient Education and Counseling.
Tsulukidze, M., Grande, S. W., & Gionfriddo, M. R. (2015). Assessing Option Grid practicability and feasibility for facilitating shared decision making: An exploratory study. Patient Education and Counseling, 98(7), 871–877.

See the full encounter lifecycle in three minutes.

From the first spoken turn to the Live Coaching Score, the Validated Observer Score, and the evidence-cited export — the demonstration walks through everything on this page, live.