StudioCast
StudioExplore

The quality layer for AI audio drama.

It listens back to its own render.

It reads your story like a listener, performs it like a full cast, and checks every line before the world hears it.

Open the studioHear a published scene
The Speckled Bandcache
verdictHold1 of 3 lines held
L3 · Helen

You know me, then?

directed surprise, heard neutral

Both paths heard it flat and agreed with each other, so the line does not ship. One path agreeing with the direction would have cleared it.

Every stage hands the next something it can check.

everything below ships

stage 01

Story Analysis

hands over a cliffhanger band and the weakest beat

  • Cliffhanger band

    A relative band with the rationale and the sources behind it, never a score.

  • One-click rewrite

    The weakest beat comes back stronger, labelled demo or live.

  • Emotional arc

    Named from six shapes, with the reference curve drawn over your own.

  • Stability band
  • Second-episode hook
  • Foreshadowing
stage 02

The Director

hands over one direction per line

  • A cast per character

    With the reasoning for each voice, shown rather than hidden.

  • Emotion and delivery

    Every line gets a directed emotion and a delivery note.

  • SFX spotted

    Cues read out of the prose and matched to a bed and SFX bank.

  • Music bed by mood
  • Ducked under dialogue
  • Voice lock
stage 03

Audio QC

hands over publish, or hold

  • Two emotion paths

    An acoustic read and an audio LLM, run independently on every line.

  • Held only when both disagree

    One path agreeing with the direction is enough to clear the line.

  • The acoustic leg

    A local SER model where the weights exist, a prosody proxy otherwise.

  • The audio judge
  • Every reading records its rung
  • Transcript diff
  • Pace
  • Loudness
  • Voice distinctiveness
  • Voice consistency
  • Pronunciation
  • Regenerate one line
after the verdict

Ship it

hands the listener a link

  • The held take against the fix

    Same line, same voice, side by side. Only the direction changed.

  • A share page

    Publish to a link that plays anywhere. No account needed.

  • A generated cover

    Typographic artwork built for the scene from its own mood.

  • MP3 export
  • A public feed
  • Appended to the flywheel

above stage 3

The Producer can overrule the rule. It cannot hear the audio.

  • Reads the measurements. Never the audio.
  • Never replaces the rule's verdict. Both are returned.
  • Its confidence is self-reported and uncalibrated.
  • Unknown stakes block. It does not guess.

around the stages

The rest of the workspace.

  • A panel, not one verdict

    Listener personas read the chapter in parallel and are allowed to disagree. Where they split is the finding: agreement tells a writer nothing they did not already know.

    Meet the panel
  • A report card, not a wall

    Emotion, pronunciation, continuity, pacing and loudness each resolve to one row with its own findings, above the Producer's release decision.

    Read the rows

measured, not claimed

The classifiers are scored, and reported as agreement.

Scored against a hand-labeled public-domain set: agreement with human labels, never a claim about absolute quality. The emotion labels are written affect rather than recorded speech, and the story signals stay relative bands with rationale.

classifieraccuracycohen kappamacro-F1
Emotionn=300.8670.8430.882
Story arcn=140.8570.8290.781

Making AI audio is easy. Trusting it is the open problem.

Paste a chapter
The quality layer for AI audio drama.Built for the Pocket FM Zero to One hackathon, IIM Bangalore.