The unglamorous half of speech AI goes Apache 2.0: Superwhisper opens a 462MB transcript cleaner — MarkTechPost
Superwhisper released S1-mini, a 596M-parameter text normalizer that punctuates raw speech-recognition output and resolves fillers and self-corrections. The quantized GGUF build is 462MB, runs on laptop CPUs, and reports 94.8% token accuracy.
Positive-sum angle: Normalisation is the step every dictation, clinical and legal transcription product rebuilds in-house. Publishing it under Apache 2.0 hands smaller builders a solved layer, so effort moves up to domain vocabulary and workflow, where local knowledge actually separates one product from another.
What's the impact: SEA founders building clinical, legal and meeting-notes tooling should stop writing normalisation code and start collecting Bahasa, Thai and Singlish correction pairs. The fine-tune is the defensible asset now, not the pipeline.