GitHub – nyrahealth/CrisperWhisper: Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection

CrisperWhisper is an advanced variant of OpenAI’s Whisper, designed for fast, precise, and verbatim speech recognition with accurate (crisp) word-level timestamps. Unlike the original Whisper, which tends to omit disfluencies and follows more of a intended transcription style, CrisperWhisper aims to transcribe every spoken word exactly as it is, including fillers, pauses, stutters and false starts.

100R — about

Hundred Rabbits is an artist collective that documents low-tech solutions with the hope of building a more resilient future. We live and work aboard a 10 m sailboat named Pino in remote parts of the world to learn more about how technology degrades beyond the shores of the western world.

Apples, Trees, and Quasimodes – System Stack

But there’s a subtle difference. For Engelbart, augmentation meant complexity: bootstrapping a system so wild it demanded co-evolution between user and tool. For Nelson, it meant endless layers of possibility. For Raskin, it often meant protection or constraint. Humane computing wasn’t only about empowerment… often it was about shielding users from mistakes, overload, and confusion.

AND

The humane thread survives, but only outside the center—in the tools that don’t have to answer to quarterly earnings, in projects that refuse to die just because they don’t fit the market. The Dormouse lineage isn’t gone. It just doesn’t live where the money is, because it can’t. If you want your computer to be humane in the deeper sense—not an appliance, but an instrument for thought—you have to look to the margins. That’s where it has always been, and where it still is today. If it survives, that’s where it’ll still be.

Methodology | Pew Research Center

Audio from video files was extracted and passed to an Audio Spectrogram Transformer model finetuned on the AudioSet dataset. This AST model inputs audio sequences, distinguishes speech from music, and then provides additional labels for the clip using a broad ontology of everyday sound types.
For videos where the AST model identified “speech” as the primary audio label, the full audio from the video was then passed to OpenAI’s whisper transcription model. For a balance of accuracy and fast processing time, we used the 769M-parameter “medium” version of this model. On English-language speech, Whisper performs speech recognition and transcription. On speech in languages other than English, the model also performs translation and returns English-language transcriptions.
All thumbnail images and slideshow images were passed through an optical character recognition (OCR) system using the python library EasyOCR. This OCR pass identified and extracted any text that could be read in the images.
All thumbnail images and slideshow images were also passed through moondream2, a lightweight vision language model that can perform text generation conditioned on an image. We used this model to produce short descriptions of the subject of each image.

MOSS: A System for Detecting Software Similarity

Moss (for a Measure Of Software Similarity) is an automatic system for determining the similarity of programs. To date, the main application of Moss has been in detecting plagiarism in programming classes. Since its development in 1994, Moss has been very effective in this role. The algorithm behind moss is a significant improvement over other cheating detection algorithms (at least, over those known to us).

How thousands of ‘overworked, underpaid’ humans train Google’s AI to seem smart | Google | The Guardian

–Funny how many times we “discover” this pattern in new spaces.

“At first they told [me]: ‘Don’t worry about time – it’s quality versus quantity,’” she said.

But before long, she was pulled up for taking too much time to complete her tasks. “I was trying to get things right and really understand and learn it, [but] was getting hounded by leaders [asking], ‘Why aren’t you getting this done? You’ve been working on this for an hour.’”

Two months later, Jackson-Artis was called into a meeting with one of her supervisors, questioned about her productivity, and was asked to “just get the numbers done” and not worry about what she’s “putting out there”, she said. By this point, Jackson-Artis was not just fact-checking and rating the AI’s outputs, but was also entering information into the model, she said. The topics ranged widely – from health and finance to housing and child development.