Diarized Transcripts
Who spoke is a fact about the recording rather than a guess about the audio. Each leg is captured on its own channel, so attribution is read off the file and nothing has to work it out acoustically.
Transcript, speaker by speaker
two channels- 0:01Marcus Hale
Adoptiv service desk, this is Marcus.
- 0:04Dana Whitlow
Hi, it is Dana Whitlow at Meridian Freight.
- 0:09Dana Whitlowsame channel, 1.4 s gap
I am calling about the invoice that went out on Tuesday.
- 0:14Marcus Hale
Let me open that. Invoice ending 4192?
- 0:18Dana Whitlow
That is the one.
- 0:21Marcus Hale
It went out at the old rate. I will reissue it today.
Product figures from the platform’s own defaults - not customer averages
How it works.
Channel beats a guessed speaker every time
Where a provider returns both a channel tag and an inferred speaker index, the channel wins. Preferring the inference over a channel that was right there is exactly the fault that swapped agent and customer on calls carrying the answer.
Speech engines are asked not to diarise at all
The four engines that can do it are asked for the channels and never for voiceprint separation. Two transcribe the channels independently with speaker labelling switched off, one sets channel diarisation with the labels fixed in channel order, and one asks for multichannel output kept apart while refusing to diarise.
A row closes on a change or a pause
Engines that answer in words rather than phrases are regrouped here: the segment holds while the channel holds and the gap stays under 1.2 seconds. Cross either and it ends, which is why one person can occupy two rows in succession.
One control fixes a call that came out backwards
Where the convention was reversed on a particular recording, a single toggle flips the labels for that call and no other. Offsets are stored with the segments, so the flip costs a redraw rather than another run.
One moment in every read.
Every read passes through the same seven. Diarized Transcripts is the lit one, and everything either side of it is a different page in this category.
- 01Source
the call or the thread it reads
- 02Transcribe
audio into words, with speakers
- 03Read
the pass over the whole of it
- 04Judge
the score, the sentiment, the intent
- 05Extract
the fields and follow-ups pulled out
- 06Write
what lands back on the record
- 07Review
a person checking the machine
The specifics.
8 facts- Convention
- Channel 0 is the agent leg and channel 1 the customer. Any other index renders under the generic label Speaker
- Segment break
- A channel change, or on word level engines the same channel pausing past 1.2 seconds
- Each segment
- Channel, text, an offset from the start, and a length where the provider gives one
- Mono audio
- Falls back to plain text with legacy prefix detection. No speaker is inferred
- Not every engine
- Two of the wired speech paths return no channel tags, and those calls take the fallback
- Searching it
- One screen has it. The voicemail list carries a transcript search running on Postgres full text; the call list searches names, agents and numbers and never the words
- Correcting a word
- Not offered. What the engine returned stands; masking strong language is a setting
- Not the same as
- This is the record of who said what. Reading it is what the analysis pass does
What it reads, and what reads it.
Nothing here invents its input. These are where the material comes from, and where the verdict goes afterwards.
More in Intelligence
12 capabilitiesAI that proposes edits to the record - an insight becomes a field once you accept it.
The model proposes record updates with its confidence and the quote it heard them in, shown as a before-and-after you accept or reject row by row.
Positive, neutral or negative on every analysed call, stored in a column of its own so the ones that went wrong are a filter rather than a listening exercise.
Chosen words rather than a fixed list for how the agent sounded, how the customer did, and the conversation overall - the manner behind the sentiment.
Communication, resolution, satisfaction and compliance scored nought to ten beside a separately judged overall, on every call that clears the analysis gates.
Compliance scored nought to ten against what the vertical says matters, because no phrase list exists to recite. A missing reading is excluded, never passed.
What the customer was actually asking for and what the agent committed to, extracted as structured intents rather than left inside the transcript.
Callbacks and meetings pulled out with their times, a confidence figure and the words they came from, ready to accept in one click or reject in one.
Collections, healthcare, insurance, admissions and eleven more - each shaping what the analysis looks for without letting anyone break the output shape.
Threads summarised and read for intent and urgency, but only when a participant matches a record, so the model is never called on mail that is not yours.
Draft from an instruction or rework a message you have, landing in the composer with a suggested subject and going nowhere at all until you send it.
Describe what you track or paste a spreadsheet's columns, and get a named field group back with types, options and placeholders to approve one at a time.
A sentence becomes a grid filter and a sort order with the reading explained back, on the toolbar of every grid that carries a column menu of its own.
The rest of the platform.
Five more categories, all on the same record and the same bill. Each card names three of its capabilities, so you can tell from here whether it is worth opening.
Intelligence
See diarized transcripts on your own floor.
Thirty minutes, your numbers and your data. We will set diarized transcripts up live and you can decide from the thing itself rather than from this page.
14-day trial · no card · migration included