Can AI hear you? Can it see you?

MIS 752 · Lab 12 Lite · Book Ch. 10 · no coding · voice and vision · every patient here is invented
In the exam room
someone speaking into a microphonethe visit is spoken
➜
a recording studio with headphonessomeone listens
➜
people reading x-rays on light boxessomeone looks closely
➜
a nursing student writing in a notebookwrites the note
➜
two people reviewing a printed documentchecks it and signs
In AI
🎙️a microphone records
➜
👂speech to text hears
➜
📷an image model looks
➜
🗂️a model drafts the note
➜
✍️a human checks and signs
A physical exam is looking and listening. AI can now do both: an ambient scribe listens to a visit and drafts the note, and an image model looks at a photo and makes a call. Neither is finished until someone checks the work, word by word and group by group.
📖 Read Chapter 10, Medical Imaging and Computer Vision (p. 215) in the course textbook ➜
📖 And for the scribes: 20.2 Ambient Clinical Intelligence: How Scribes Become Agents (p. 468)
Same book on Canvas: course files
🔬 What is real here. The visit script, every mistake Whisper made, and every skin-audit number come from the Lab 12 notebooks' own runs on a free Colab GPU (September 13, 2026). The model you pick below runs live. Simulated: the busy-corridor transcript is assembled from those real Whisper mistakes, and the two audio clips use a Mac voice. The patient is invented, and no photo of anyone's skin appears on this page.
Photos: Josh Hawkins and Becca Schwartz, UNLV. Real UNLV spaces; the patient and the visit in this lab are invented.

0 · Connect a model

Paste the free OpenRouter key from Lab 1. It stays in this browser tab only: it is never saved and never sent anywhere except OpenRouter.

1 · Four ideas before you start

Open each card, make your guess, then read the answer. Every number is from the notebook's real run.

Whisper wrote down the visit with 5 mistakes in 237 words: a 2.1% word error rate. Is it ready for the chart?

Not on that number alone. Two of the five mistakes were "A1c" (heard as "IAC") and "lisinopril" (heard as "lisenalprol"): the diabetes test and the blood-pressure drug. Word error rate counts "the" and "lisinopril" the same. A clinician does not.

The doctor says, "Your blood pressure today is 148 over 92." Which SOAP drawer does it go in?

OBJECTIVE: the clinic measured it. SUBJECTIVE is what the patient reports ("my feet tingle at night"), ASSESSMENT is the clinician's judgment, PLAN is what happens next. The notebook's real note filed 148/92 under both Subjective and Objective.

Whisper heard "lisenalprol". The note says "lisinopril", spelled right. What kind of fact is that?

A lucky fill-in: it was said, Whisper missed it, and the model wrote it anyway. Right this time, with nothing to go on. The other two kinds: invented (nobody said it, like the "type 2" the real note added) and carried mishearing (Whisper's mistake copied into the note).

Can a word error rate be higher than 100%?

Yes, because words that nobody said count as errors too. With a robotic voice and loud noise, Whisper wrote 288 words nobody said: a word error rate of 158.6% in the notebook's robot-voice run. Then, at even louder noise, it scored "better" because it gave up and wrote less.

7 · Try it live

Three free pages, no account. One lets you hear what a speech model hears; two show how fairness audits work.

Whisper Web on Hugging Face: transcribe from a URL, from a file, or by recording
Whisper Web, a free Hugging Face Space by Xenova. It runs in your browser; no account needed.

🤗 Try it live on Hugging Face: can a real speech model hear you?

Whisper Web runs OpenAI's Whisper speech model inside your own browser. The first load downloads the model, so give it a minute.
  1. Download this page's two invented clips: 🔇 quiet room and 📢 busy corridor.
  2. In Whisper Web choose From file and transcribe each clip. Or choose Record and read an invented line, such as "Increase the lisinopril from ten to twenty milligrams, and stop the glipizide."
  3. Compare the two transcripts word by word. Use only invented dialogue or these two clips: never a real patient's voice.
Open Whisper Web on Hugging Face ➜
Which words broke first in the corridor clip, and were they the ones a clinician cares about most?

Usually the drug names break first: they are rare words the model heard less often in training, while common words like "increase" and "milligrams" survive. In the notebook's own noisy run, lisinopril came out as "lisen awful" and glipizide as "glitz eye", while "10 to 20 milligrams" survived. Numbers often survive, but "ten to twenty" heard as "two" would be a tenfold dosing error. So the words a clinician cares about most are often the hardest to hear, which is why a word error rate alone says too little.

🌐 See it live: a skin-tone scale made for fairness testing

Google Research published the 10-shade Monk Skin Tone Scale, created by sociologist Dr. Ellis Monk, for checking whether technology works across skin tones. This lab's audit uses the older 6-type Fitzpatrick scale.
  1. Open the page and look at the 10 shades.
  2. Think of this lab's audit: Fitzpatrick puts all the darkest skin into types V and VI, the group with only 11 photos. How many of the 10 shades sit at that dark end?
  3. Read why the scale was made, and what it is used for.
Open the Monk Skin Tone Scale ➜

🌐 See a real audit: Gender Shades (MIT Media Lab)

Joy Buolamwini's project audited commercial face-analysis systems on 1,270 faces, split by skin type and gender together.
  1. Open the project page.
  2. Find how the faces were grouped (darker women, darker men, lighter women, lighter men).
  3. Find the worst and the best error rates.
Open Gender Shades ➜

8 · Design your own audit

This is the real skill. Pick a listening or looking tool you would want in your own work, then plan the check that must happen before anyone trusts it. No code: plain English.

📖 Read more in the book: 20.3 Patient Consent for Ambient Recording (p. 469)📖 Read more in the book: 23.1 Sources of Bias: Historical, Label, and Selection Bias (p. 546)📖 Read more in the book: 10.2.3 Dermatology (p. 223)

🎙️ Now try it on a visit you write

Write a short invented visit, mark the facts that must survive, and let the model draft the note with the same SOAP rules.

9 · Hand it in (Canvas, Lab 12)

1. Download your submission with the button below, then upload the file to the Lab 12 assignment on Canvas. It holds every run, your predictions, both sides of each example, your audit plan, your own visit, and everything you wrote.

2. Answer these five, a few sentences each. Each asks why:
  1. When your prediction was wrong, what had you assumed about how AI listens or looks that turned out not to be true?
  2. The busy corridor. Pick one fact that changed between the two notes. Follow it downstream: who acts on it next, and what happens to the patient?
  3. Copy, fix, or flag. Which behavior do you want from a scribe in your clinic, and why is a correct quiet fix still a risk?
  4. The skin audit. The darkest skin types scored highest, on five cancers. Write the sentence you would put on the product label, and explain why.
  5. Your own audit. Which groups did you choose to split by, why, and what result would make you say NOT ENOUGH EVIDENCE instead of DEPLOY?

Nothing you type is stored anywhere. Download your file before you close the tab.