Files
win-dictate/win-dictation-architecture-engineering-review.md

25 KiB
Raw Permalink Blame History

Win Dictation — Architecture & Engineering Review

Architecture · Code Analysis · Engineering Verdict

A guided walk through a native Win32 C++ speech-to-text app — how it's built, why it's built that way, and an honest assessment of the code for anyone picking it up for the first time.

~3,400 lines C++ (app) Win32 + GDI+ + SDL2 + whisper.cpp Single .exe, no runtime deps Target 2-core i5, CPU-only

Contents

# Section
01 The verdict, up front
02 Stack at a glance
03 Architecture & data flow
04 The defining decision
05 Threading model
06 Source map, file by file
07 Deep dive: the UI engine
08 Deep dive: the progress estimator
09 Deep dive: the transcriber
10 History, downloads & persistence
11 Anatomy of one dictation
12 What's done well
13 What holds it back
14 Recommendations
15 Scorecard & final word

01 The verdict, up front

Bottom line

A genuinely strong, characterful single-purpose tool that punches well above hobby grade.

The hard engineering — the architecture choice, the self-calibrating progress estimator, the seam-free single-surface renderer — is thoughtful and well-executed. The app does exactly one thing and does it well on hardware most tools would choke on.

What holds it back is organizational debt, not algorithmic weakness: a 1,500-line main.cpp, a layer of vestigial child-window controls left over from an earlier design, and a pile of stale documentation that describes a GPU-streaming app this no longer is. None of it breaks the running product — but all of it raises the cost of the next person walking in.

The sections below back up every part of that judgement with specifics from the source.


02 Stack at a glance

Deliberately lean. No UI framework, no managed runtime, no garbage collector — just the OS and three libraries.

Layer Choice
Language C++ (MSVC, Release /O2 /GL /LTCG, AVX2/FMA/F16C)
UI Raw Win32 + GDI+ (immediate-mode painted surface)
Audio capture SDL2 (16 kHz mono, F32)
Inference whisper.cpp (whisper_full, CPU backend)
Networking WinHTTP (model downloads, system proxy aware)
Persistence Plain INI + UTF-8 text files (no database)
Build CMake (+ a PowerShell convenience script)
Footprint One .exe + a few DLLs + model (tiny RAM, no install)

The whole product is native code with no framework abstraction between it and the Win32 API. That's the source of both its biggest strength (a tiny, fast, dependency-light binary) and its biggest cost (everything is hand-rolled, including the widgets).


03 Architecture & data flow

A linear pipeline with a single inference step. Audio in, text out, no streaming loop.

┌────────┐   ┌────────┐   ┌────────┐   ┌────────────┐   ┌──────────┐
│ Hotkey │→  │  SDL2  │→  │PCM in  │→  │whisper_full│→  │ Insert + │
│ trigger│   │   mic  │   │  RAM   │   │   worker   │   │   paste  │
└────────┘   └────────┘   └────────┘   └────────────┘   └──────────┘
  • Hotkey — Global RegisterHotKey. Captures the previously focused window.
  • SDL2 mic — 16 kHz mono into an in-memory buffer. Near-zero CPU.
  • PCM in RAM — Accumulated under a mutex; RMS energy tracked for the meter.
  • whisper_full — One pass on a worker thread when you stop. Trim → transcribe → clean.
  • Insert + paste — Text to caret, to clipboard, into the prior window.

Two side channels run alongside the main pipeline: a progress estimator that predicts and smooths the transcription countdown, and a timing model that learns this machine's speed and feeds back into the next prediction. Results return to the UI thread exclusively via PostMessage; shared flags are std::atomic.

Mental model

Think of it as a tape recorder with a transcription button, not a live captioner. The architecture has no per-frame transcription loop at all — which, on a 2-core CPU, is the whole point.


04 The defining decision

The single most important thing to understand: this app was rebuilt from a live-streaming design into a push-to-talk batch design — and that was the right call.

An earlier version used the classic whisper.cpp streaming approach: a rolling 56 second window re-transcribed every ~0.4 seconds. That technique assumes spare cores. On the target machine — an Intel i5-7th-gen with two physical cores — the buffer backlogged, audio was re-transcribed, and the UI starved. The symptom looked like "the model is slow"; the real cause was an architecture that needed hardware the target didn't have.

The fix wasn't a faster model. It was removing the streaming loop entirely:

Aspect Old: streaming window New: push-to-talk batch
CPU while speaking Pinned — constant re-inference Near idle — just buffering
Inference calls Many per second Exactly one, on stop
Accuracy Lower — partial context windows Higher — full clip, full context
UI responsiveness Starved under load Free until the single pass
Predictability Variable lag Predictable few-second wait

This is textbook root-cause engineering: the team correctly diagnosed that the bottleneck was the shape of the work, not its size, and changed the shape. Everything else in the codebase — the batch worker, the progress estimator, the physical-core thread default — follows from this one decision.


05 Threading model

Four threads, one rule: only the UI thread touches the UI. Everything else reports back by message.

Thread Lifetime Job
UI thread Whole app Message loop, all painting, the 16 ms animation timer and 50 ms update timer.
Model preload Detached, once Loads the Whisper context off the UI thread at startup so the window appears instantly.
Transcribe worker Per clip Runs whisper_full; posts WM_APP_PROGRESS during and WM_APP_RESULT when done.
Downloader Per download WinHTTP fetch on its own thread; posts WM_APP_DLPROGRESS.

Cross-thread state is handled with discipline rather than locks where possible: std::atomic booleans (m_recording, m_busy, m_abort, g_modelLoaded, g_modelOk) gate state transitions, and the only shared buffer — the captured PCM — is protected by a dedicated mutex. The audio callback (driven by SDL's own thread) appends under that mutex; stop_and_transcribe swaps the buffer out under the same lock before handing it to the worker. That swap-not-copy handoff is a nice touch.

// transcriber.cpp — physical cores, not logical, by design
int Transcriber::default_threads() {
    unsigned hc = std::thread::hardware_concurrency();
    if (hc <= 2) return (int)std::max(1u, hc);
    return (int)(hc / 2);   // 4 logical → 2 worker threads
}

Defaulting to physical cores rather than hardware_concurrency() is the correct choice for compute-bound SIMD inference — hyperthreads contend for the same execution units and would only add scheduling overhead.


06 Source map, file by file

The app is small and mostly header-only outside the two big translation units. Here's where everything lives.

File Role Notes
main.cpp Window, painting, interaction, settings view, clipboard, paste, model selection, popups ~1,500 lines. The monolith — see §13.
transcriber.{h,cpp} SDL capture, Whisper preload/inference, progress & abort callbacks Clean, well-scoped class. The model layer.
timing.h Per-model least-squares timing model + live progress estimator + INI persistence The standout module. See §08.
history.h Session text files, UTF-8 r/w with BOM, index, pruning to 100 Self-contained, header-only.
downloader.h WinHTTP model downloader on a background thread .part + atomic rename, cancel, proxy-aware.
stats.h Lifetime usage totals + derived figures (wpm, real-time factor, time saved) INI-backed, header-only.
settings.h App settings read/write via GetPrivateProfile* Simple and transparent.
text_util.h · logging.h Transcript concatenation; timestamped file log Tiny helpers.
tests/test_core.cpp Unit checks: append logic, bad-model handling, real-WAV transcription, progress monotonicity Modest but meaningful. Built as test-core.

The decision to make most subsystems header-only and independent (timing.h, stats.h, history.h, downloader.h, settings.h) is a good one for a project this size: each is cohesive, individually readable, and free of cross-dependencies. The contrast with main.cpp — which absorbs everything else — is stark.


07 Deep dive: the UI engine

There are no buttons. Everything you see is painted onto one double-buffered surface — and that's a deliberate fix, not a shortcut.

The previous UI composited around nine separate themed child windows (buttons, statics, an edit). That produced thin hairline seams around every control — the hard-edged holes WS_CLIPCHILDREN punches per child, plus the edit's themed border. Rather than chase pixel borders, the rebuild eliminated the cause: collapse the controls into one painted region.

The model is a small immediate-mode system:

  • A flat Widget g_w[] array — each entry is a kind, a rectangle, and hover/pressed/anim state. No HWNDs.
  • LayoutWidgets() positions them; PaintSurface() draws each into an off-screen DC with GDI+, then blits once (no flicker).
  • HitTest() maps a click point to a widget; OnClick() dispatches the action.
  • An animation clock eases each widget's anim toward a target (hover 0.6, active 1.0) on a 16 ms timer that stops itself when nothing is moving — no idle CPU burn.

Two "views" — Main and Settings — render onto the same surface, toggled by SwitchView(). Dropdowns (mic, history) are the one exception: they're real top-level WS_POPUP windows, because a surface-painted dropdown would render behind the transcript edit (a child HWND always paints above its parent's surface). That's a correct, well-reasoned exception.

A sign of maturity

The popup code carries a comment never to open a MessageBox from inside it — because WA_INACTIVE self-destroys the popup mid-handler, causing a use-after-free. Recognising that class of Win32 lifetime bug, and the GDI+ "GetHDC locks the Graphics object" trap documented elsewhere, shows real depth.

The one real child window that survives is the transcript EDIT — kept because a hand-rolled text editor with selection, scrolling, IME and undo is genuinely not worth rebuilding. Pragmatic.

The cost of this approach

Custom-painted controls are invisible to screen readers and UI Automation, and the app explicitly hides focus rectangles. The transcript box is accessible; the buttons are not. For a personal productivity tool this is a defensible trade, but it's the kind of thing worth stating out loud.


08 Deep dive: the progress estimator

The crown jewel. Most apps fake a progress bar; this one runs a small statistical model that learns your machine.

Whisper only reports coarse progress (per 30-second chunk), so a naive bar jumps — 0, 34, 72, 100 — and a naive "time left" computed from stale percentages actually counts up. timing.h solves both. It has two parts.

1 — A learned timing model

Processing time is modelled as a linear function of audio length, proc = a + b·audio, fitted by decayed online least-squares. Each completed transcription feeds back a real sample; older samples decay (factor 0.97) so the model tracks the current machine state. Defaults are seeded per model family (tiny/base/small) so even the very first clip has a sane estimate, and the accumulators persist per-model in the INI.

void add_sample(double audio_sec, double proc_sec) {
    const double decay = 0.97;
    n*=decay; sx*=decay; sy*=decay; sxx*=decay; sxy*=decay;
    n+=1; sx+=audio_sec; sy+=proc_sec;
    sxx+=audio_sec*audio_sec; sxy+=audio_sec*proc_sec;
    recompute();   // closed-form slope/intercept
}

2 — A live estimator that only counts down

On stop, begin(predict) seeds a predicted total. As Whisper reports chunk progress, on_whisper() folds it in as a measurement via an EMA (α = 0.5) — nudging the estimate without the jumpy jumps. Meanwhile tick() advances a displayed "remaining" value that is strictly monotonic downward, with a clamped catch-up rate so it can speed up but never lurch backward, easing to 95% and snapping to 100% only on the real result.

void tick(double dt, float& out_frac, float& out_remaining) {
    t += dt; disp_rem -= dt;
    double raw_rem = std::max(0.0, T_hat - t);
    double err = raw_rem - disp_rem;
    if (err < 0) disp_rem += std::max(err, -maxCatchUp*dt);  // catch up, never jump back
    double frac = t/(t+disp_rem);
    if (frac > 0.95) frac = 0.95;   // park at 95% until done
    out_frac = (float)frac; out_remaining = (float)disp_rem;
}

This is more thought than most commercial apps put into a progress bar, and the test suite even asserts the progress is non-decreasing. It's the clearest signal in the codebase that someone cared about the feel of the product, not just its function.


09 Deep dive: the transcriber

The cleanest class in the project — a tidy boundary between the OS/model and the rest of the app.

Transcriber owns the Whisper context and the SDL device, and exposes a small, sensible surface: preload, reload, start_recording, stop_and_transcribe, cancel, plus state queries and two callbacks (result, progress). Inference parameters are configured sensibly for dictation — greedy sampling, no timestamps, no prior context, blank/non-speech suppression, temperature 0 — and an abort_callback lets a long transcription be cancelled mid-flight.

Two small details worth calling out:

  • Silence trimming. Before inference, leading/trailing silence is trimmed from the clip — cheaper and more accurate than transcribing dead air. (Note: this trims the buffer; it does not auto-stop recording — see the doc-drift note in §13.)
  • Output cleanup. clean_text() strips Whisper's [BLANK_AUDIO] / [NOISE] artifacts and trims whitespace, so the user never sees model noise.

The class is also defensively coded: a missing model file makes preload return false cleanly (the test suite verifies this), start_recording bails if a device won't open, and clips under ~0.3 s short-circuit to an empty result rather than invoking the model.


10 History, downloads & persistence

No database, no registry sprawl — everything is a file next to the executable. Transparent and portable.

Session history

Each session is one UTF-8 text file in history\, named by timestamp. The clever bit is live archiving: the first clip of a session creates the file; subsequent clips rewrite the same file with the full text. So a session is always one tidy, crash-safe file — not a scatter of fragments — and it appears in the History popup immediately. A g_sessionPath global plus a FinalizeSession() helper handle the edge cases (manual edits, typed-only sessions, loading an old entry without resurrecting it). The list is capped at 100 with automatic pruning.

A real bug was fixed here

An earlier version stamped every fresh transcription as "loaded from history," so the duplicate-guard silently skipped archiving — sessions never reached disk while the UI claimed "Saved." The fix (live archiving + a corrected guard) is documented and shows the team chasing subtle state bugs to ground.

Model downloads

The downloader is more robust than it needed to be, in a good way: it streams to a .part file then does an atomic rename on success (no half-files), honors the system proxy, follows the Hugging Face → CDN redirects, supports cancellation, allows only one download at a time, and sweeps up stray .part files at startup.

Settings, timing & stats

All three live in a single win-dictation.ini under different sections — app settings, per-model timing accumulators, and lifetime stats. Using the OS's own GetPrivateProfile* API means zero parsing code and a file a user can read and edit by hand. For an app of this scope, that's exactly the right level of machinery.


11 Anatomy of one dictation

Following a single clip end-to-end ties the whole system together.

  1. WM_HOTKEY fires → the app records g_prevForeground and the current selection (EM_GETSEL) so it knows where to paste and where to insert.
  2. g_tx.start_recording() opens the SDL device; the audio callback appends PCM under the capture mutex and updates the RMS energy meter.
  3. The 50 ms UI timer animates the level meter and ticks the on-screen recording clock.
  4. Second WM_HOTKEYg_est.begin(g_timing.predict(len)) seeds the progress estimate; g_tx.stop_and_transcribe() swaps the buffer to a worker thread.
  5. The worker runs run_inference()whisper_full. Whisper's progress callback posts WM_APP_PROGRESS; the estimator's on_whisper() EMA-folds it in.
  6. On completion the worker posts WM_APP_RESULT.
  7. The UI thread then, in order: snaps progress to 100%, records a real timing sample (add_sample + SaveTiming), updates lifetime stats, inserts the text at the saved caret with smart spacing via EM_REPLACESEL (undoable), archives the session, copies to the clipboard, and pastes into g_prevForeground.

Every piece of the architecture shows up in that one trip: the atomics, the message hand-back, the estimator, the learned timing feedback loop, the editable transcript, the live history. It's a coherent design.


12 What's done well

  • The architecture fits the hardware — Push-to-talk batch over streaming is the correct response to a 2-core CPU, reached by genuine root-cause analysis rather than knob-twiddling.

  • The progress estimator is exceptional — Decayed online least-squares + EMA fusion + a strictly monotonic countdown is far beyond what the task demanded — and it shows in the feel.

  • Seam-free UI by elimination, not patching — Collapsing nine child windows into one painted, double-buffered, DPI-aware, self-throttling surface removed the problem at its source.

  • Clean module boundaries (outside main) — Header-only, dependency-free subsystems (timing, history, downloader, stats, settings) are each individually readable and testable.

  • Robustness in the right places — Atomic rename downloads, crash-safe live history, graceful missing-model handling, cancellable inference, single-instance mutex, model preload off the UI thread.

  • Real product thoughtfulness — Smart insertion spacing, undoable edits, auto-paste into the prior window, auto-hide, learned timing, friendly stats. These are details a careful builder adds.


13 What holds it back

All fixable, and none of it affects the running app. But it's exactly what a newcomer trips over.

main.cpp is a 1,500-line god object

UI, layout, painting, the entire settings screen, clipboard, paste mechanics, model selection, the popup window class, and stats formatting all live in one translation unit with dozens of globals. It works, but it's the hardest part of the codebase to onboard into. Splitting the settings view, the popup, and the painting helpers into their own files would pay for itself quickly.

⚠ Vestigial child windows & duplicate code paths

Startup still creates ~9 owner-draw child controls (record, pin, copy, paste, clear, two selects, two statics) and then immediately hides all but the transcript edit. Their dead WM_COMMAND handlers duplicate the painted-widget OnClick logic — e.g. the Copy action exists in two near-identical places. Leftovers from the rebuild that should be deleted.

main.cpp — CreateWindow(...) blocks then ShowWindow(..., SW_HIDE)

Documentation describes a different app

This is the most actively misleading issue. CUDA-SETUP.md, QUICK-REBUILD-GPU.md and FIXES-APPLIED.md describe a streaming, VAD, ring-buffer, 24-thread, RTX 3090 design that no longer exists. build.ps1 still hunts for CUDA and downloads base.en though the product is a CPU-only tiny.en app. TESTING.md references a test-audio.exe the CMake doesn't build (it builds test-core). A newcomer reading the docs would form a completely wrong mental model.

The README claims a feature that isn't there

Both README.md and CHANGES.md describe a "500 ms silence auto-end timer." The recording loop has no such logic — it only auto-stops at the 10-minute safety cap. (Silence is trimmed before inference, which is likely the source of the confusion.) Either implement it or remove the claim.

main.cpp WM_TIMER recording branch vs README "Audio Processing"

⚠ Heavy reliance on global mutable state

The UI is coordinated through dozens of file-scope globals (g_*), a mix of atomics and plain values. It's manageable at this size and the threading is disciplined, but it makes the code hard to reason about in isolation and easy to break with a careless edit.

⚠ Build declares C++11 but uses C++17

CMakeLists.txt sets CMAKE_CXX_STANDARD 11, yet the code uses std::size() (C++17). It compiles only because MSVC's default is newer. Set the standard to 17 explicitly so the build is honest and portable.

⚠ Minor: redundant color systems & no in-app hotkey editor

Three overlapping palettes coexist (CR_* COLORREF, T_* GDI+ Color, C_* aliases). And changing the hotkey requires hand-editing the INI — a natural gap given the polished Settings screen already exists.


14 Recommendations

If the next session had a short to-do list, this would be it — ordered by payoff for effort.

# Task Priority
1 Purge or archive the stale docs — Delete or clearly mark CUDA-SETUP.md, QUICK-REBUILD-GPU.md, FIXES-APPLIED.md, TESTING.md and DESIGN.md as describing the retired streaming design. This is the single biggest improvement to onboarding, and it's nearly free. High
2 Reconcile the README with reality — Remove the "500 ms silence auto-end" claim (or implement it). Update the model table and build commands to match the CPU-only product. High
3 Delete the vestigial child windows — Remove the hidden owner-draw controls and their dead WM_COMMAND handlers so there's exactly one code path per action. De-duplicate Copy. Medium
4 Break up main.cpp — Lift the Settings view, the popup window, and the GDI+ drawing helpers into their own files. Even a mechanical split dramatically improves navigability. Medium
5 Fix the build standard & align build.ps1 — Set CMAKE_CXX_STANDARD 17. Strip the CUDA detection from the build script and default it to fetching tiny.en. Medium
6 Add an in-app hotkey picker — The Settings surface already exists; surfacing the hotkey there closes an obvious UX gap and removes a troubleshooting step. Low

15 Scorecard & final word

Category Score
Architecture & design █████████░░ 9.2
Performance fit for target █████████░░ 9.3
UX & polish ████████░░░ 8.7
Robustness & error handling ███████░░░░ 7.5
Code organization █████░░░░░░ 5.5
Maintainability █████░░░░░░ 5.8
Testing ████░░░░░░░ 4.8
Documentation accuracy ███░░░░░░░░ 3.8

Overall: 7.1 / 10 — Strong, with cleanup debt.

An impressive core wrapped in organizational and documentation drift. The engineering earns a high mark; the housekeeping pulls the average down.


Final word

Win Dictation is a good codebase — at its core, an impressive one. The architectural judgement (batch over streaming), the standout progress estimator, and the seam-free renderer are the work of someone who diagnoses root causes and cares about how software feels. Those are the hard parts, and they're done right.

What separates it from "great" is entirely recoverable: a monolithic main file, dead code from a prior design, and documentation that actively describes a different application. A focused day of cleanup — most of it deletion — would lift this from "strong for its niche" to "exemplary small-app code." The good news for anyone inheriting it: the bones are excellent, and the to-do list is short.


Win Dictation — Architecture & Engineering Review Native Win32 C++ · GDI+ · SDL2 · whisper.cpp · CPU-only · target Intel i5-7th-gen (2C/4T) Assessment based on a full read of the current source tree.