A lightweight, private speech-to-text application powered by moondream/parakeet-redux. Automatically converts uploaded video or audio files into timestamped text transcripts locally on your computer at 113× real-time speed.
| Step 1: Input | Step 2: AI Action | Step 3: Result |
|---|---|---|
| Media Upload Upload any video ( .mp4, .mov) or audio (.mp3, .wav) file. |
1.58-Bit AI Processing System FFmpeg standardizes audio; Parakeet Redux transcribes instantly on CPU. |
Timestamped Transcript View sentence segments with time codes and download .md output. |
Before running the application, make sure your computer has Python 3.9 or higher and system-level FFmpeg installed to extract audio from video containers.
- Windows (PowerShell Admin / Chocolatey or Winget):
winget install --id FFmpeg-FFmpeg.FFmpeg -e
- Ubuntu / Debian:
sudo apt update && sudo apt install ffmpeg - macOS:
brew install ffmpeg
Run the following single PowerShell command to install all required dependencies:
pip install streamlit moondream numpyLaunch the Streamlit web interface with the following command:
streamlit run app.pyThe web interface will open automatically in your browser at http://localhost:8501. On the first upload, the application downloads the lightweight 178 MB model weights directly into your local cache.
├── app.py
├── requirements.txt
└── README.md
- 🎙️ Local Meeting & Podcast Transcription: Convert multi-hour meeting recordings into text locally with zero cloud API costs or privacy leaks.
- 🔍 Voice Note Search Vault: Turn scattered voice memos into searchable text files for personal knowledge management.
- 📚 Lecture Chapter Indexing: Automatically split long university lectures into readable notes using built-in pause detection.
- 🌐 Global Accent & Language Testing: Test transcription accuracy across 25 supported global languages natively.
- 💻 Hands-Free Developer Voice Notes: Speak technical thoughts or bug descriptions directly into local structured markdown logs.
- ⏱️ Word-Level Precision Timestamps: Highlight exact spoken words in real time as media plays back.
- 🏷️ Speaker Identification (Diarization): Automatically label different speakers in multi-person meetings.
- 📊 Interactive Transcript Search & Filter: Instantly filter transcript lines by keyword or timestamp range.
- 📝 Automated AI Meeting Summarization: Feed generated transcripts into local LLMs to produce executive summaries.
- 📂 Batch Folder Processing: Drag and drop an entire directory of audio files for background processing.
moondream parakeet-redux 1.58-bit-ai ternary-quantization local-speech-to-text ffmpeg-audio-extraction streamlit-transcriber cpu-speech-recognition faster-than-whisper