Files
win-dictate/release/WinDictation

# Win Dictation — Voice to Text for Windows

![screenshot](screenshot.png)

A push-to-talk speech-to-text utility for Windows. Press a hotkey, speak, and your words land in whatever app you were just using — fully offline, powered by Whisper.

## Quick Download

**[Download Latest Release (WinDictation.zip)](release/WinDictation.zip)**

Extract the ZIP file and run `win-dictation.exe`. The release includes all required DLLs and a Whisper model.

For new releases, run `.\src\package.ps1` after building, then attach `dist\WinDictation.zip` to a [Gitea Release](https://docs.gitea.com/usage/packages/overview) on your instance.

---

## Documentation

- **[User Manual](win-dictation-user-manual.html)** — How to use every feature
- **[Architecture & Engineering Review](win-dictation-architecture-engineering-review.html)** — Deep dive into the codebase for developers

---

## Features

### Performance
- **Physical-core threading**: Uses one thread per physical core for efficient batch transcription
- **CPU-only**: Optimised for ordinary laptops — no GPU required
- **Self-calibrating progress**: Learns transcription speed per model and machine, delivering a smooth countdown

### User Interface
- **Single-surface rendering**: No inter-window seams or hairlines — the entire UI is one painted surface
- **Dark theme**: Calm, elevated card design with hover/press feedback
- **Per-monitor DPI awareness**: Looks sharp at any display scale
- **System tray**: Minimise to tray, global hotkey to record

### Audio Processing
- **Push-to-talk**: Press Ctrl+Shift+Space, speak, press again to transcribe
- **Multiple microphones**: Select from all available input devices
- **16 kHz sample rate**: Optimised for Whisper
- **Silence trimming**: Leading and trailing silence is trimmed before transcription

## Quick Start

1. Launch `win-dictation.exe`
2. Click into wherever you want text, then press `Ctrl+Shift+Space`
3. Speak, then press `Ctrl+Shift+Space` again
4. Your words appear in the app you were using

Full instructions in the [User Manual](win-dictation-user-manual.html).

### Controls
- **Record/Stop**: Click the pill or press `Ctrl+Shift+Space`
- **Hide window**: `Ctrl+Shift+H`
- **Pin**: Keep window always-on-top
- **Copy / Paste / Clear**: Text actions below the transcript
- **Model / Mic / History**: Select from popup menus

## Building from Source

### Prerequisites

- **Windows 10/11**
- **CMake** 3.5+
- **Visual Studio 2022** with C++ workload
- **SDL2** (included in deps/)

### Build Steps

```powershell
cmake -S . -B build -G "Visual Studio 17 2022" ^
  -DSDL2_DIR="deps/SDL2-2.28.5/cmake"
cmake --build build --config Release
```

The executable will be at `build\bin\Release\win-dictation.exe`.

## Technical Details

### Architecture

```
Hotkey → SDL Capture → Stop → Batch whisper_full → Text → Auto-paste
```

1. **SDL audio capture**: 16 kHz mono recording into memory. CPU stays near idle while recording.
2. **Stop-and-transcribe**: One `whisper_full` call processes the full clip at once.
3. **Progress estimation**: Decayed online least-squares model per machine/model, fused with whisper's chunk progress into a strictly monotonic countdown.
4. **Text output**: Inserted at cursor with smart spacing, copied to clipboard, and optionally auto-pasted into the window you came from.

### Models

Place `.bin` files in `models/` next to the executable, or download them from the Settings screen in-app:

| Model | Size | Best for |
|-------|------|----------|
| tiny.en | ~75 MB | Fastest — everyday dictation |
| tiny.en-q8_0 | ~42 MB | Same speed, smaller file |
| base.en-q5_1 | ~59 MB | Good accuracy bump for little cost |
| base.en | ~142 MB | More accurate; still reasonable on two cores |
| small.en-q5_1 | ~182 MB | Accurate, but slower |
| small.en | ~466 MB | Most accurate — and slowest |

## Project Structure

```
win-dictation/
├── src/                # Application source
│   ├── main.cpp        # Window, painting, interaction, settings, clipboard, popups
│   ├── transcriber.*   # SDL capture, Whisper preload/inference, progress callbacks
│   ├── timing.h        # Learned timing model + live progress estimator
│   ├── history.h       # Session text files, index, pruning
│   ├── downloader.h    # WinHTTP model downloader (background thread)
│   ├── stats.h         # Lifetime usage totals + derived figures
│   ├── settings.h      # INI persistence
│   ├── text_util.h     # Transcript concatenation helpers
│   ├── logging.h       # Timestamped file log
│   └── tests/          # Unit tests (test-core.exe)
├── whisper/            # Whisper.cpp library
├── ggml/               # GGML tensor library
├── models/             # Whisper model files
├── history/            # Saved dictation sessions
├── release/            # Pre-built package
└── CMakeLists.txt      # Build configuration
```

## License

[GNU General Public License v3.0](LICENSE)

This project is licensed under GPLv3. It incorporates [whisper.cpp](https://github.com/ggerganov/whisper.cpp) (MIT), [ggml](https://github.com/ggerganov/ggml) (MIT), and [SDL2](https://www.libsdl.org/) (zlib).

## Resources

- [User Manual](win-dictation-user-manual.html)
- [Architecture & Engineering Review](win-dictation-architecture-engineering-review.html)
- [Whisper.cpp](https://github.com/ggerganov/whisper.cpp)
- [Model Downloads](https://huggingface.co/ggerganov/whisper.cpp)