Vibecode Superwhisper
track this build5 steps, step by step0%A push-to-talk recorder that transcribes and pastes text into the active app is very buildable for one person, especially on macOS.
You are building a lean indie version of Superwhisper. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # Superwhisper indie build ## Goal Build the smallest trustworthy replacement for the core Superwhisper workflow for one developer or a tiny team. ## Scope Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - beautiful native UX - presets - vocabulary/profile tuning - app-wide polish - support If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements: - A global hotkey (default: hold right Option) records my mic while held, stops on release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to ship. - Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp (small.en by default, model path in a config file). Works fully offline, no cloud speech APIs. - Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or osascript, then restore my previous clipboard). - Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in .env) to fix punctuation and drop filler words, then paste. If no key is set, this mode just does a plain paste. - Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10 transcripts with copy buttons. - Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the audio after transcription. No accounts, no telemetry. - Out of scope: per-app presets and custom vocabulary tuning. One good general mode. - README: mic + accessibility permissions to grant, how to download the whisper model, and a note that the first run will trigger macOS permission prompts. ## Required capabilities - macOS automation/accessibility permissions - speech API or local Whisper - optional LLM API - hotkey library ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of Superwhisper. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # Superwhisper indie build ## Goal Build the smallest trustworthy replacement for the core Superwhisper workflow for one developer or a tiny team. ## Scope Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - beautiful native UX - presets - vocabulary/profile tuning - app-wide polish - support If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements: - A global hotkey (default: hold right Option) records my mic while held, stops on release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to ship. - Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp (small.en by default, model path in a config file). Works fully offline, no cloud speech APIs. - Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or osascript, then restore my previous clipboard). - Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in .env) to fix punctuation and drop filler words, then paste. If no key is set, this mode just does a plain paste. - Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10 transcripts with copy buttons. - Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the audio after transcription. No accounts, no telemetry. - Out of scope: per-app presets and custom vocabulary tuning. One good general mode. - README: mic + accessibility permissions to grant, how to download the whisper model, and a note that the first run will trigger macOS permission prompts. ## Required capabilities - macOS automation/accessibility permissions - speech API or local Whisper - optional LLM API - hotkey library ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of Superwhisper. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # Superwhisper product brief ## Problem A push-to-talk recorder that transcribes and pastes text into the active app is very buildable for one person, especially on macOS. ## Product outcome Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - macOS automation/accessibility permissions - speech API or local Whisper - optional LLM API - hotkey library ## Explicit non-goals for v1 - beautiful native UX - presets - vocabulary/profile tuning - app-wide polish - support ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements: - A global hotkey (default: hold right Option) records my mic while held, stops on release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to ship. - Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp (small.en by default, model path in a config file). Works fully offline, no cloud speech APIs. - Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or osascript, then restore my previous clipboard). - Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in .env) to fix punctuation and drop filler words, then paste. If no key is set, this mode just does a plain paste. - Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10 transcripts with copy buttons. - Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the audio after transcription. No accounts, no telemetry. - Out of scope: per-app presets and custom vocabulary tuning. One good general mode. - README: mic + accessibility permissions to grant, how to download the whisper model, and a note that the first run will trigger macOS permission prompts. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted Superwhisper capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# Superwhisper indie build ## Goal Build the smallest trustworthy replacement for the core Superwhisper workflow for one developer or a tiny team. ## Scope Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - beautiful native UX - presets - vocabulary/profile tuning - app-wide polish - support If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements: - A global hotkey (default: hold right Option) records my mic while held, stops on release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to ship. - Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp (small.en by default, model path in a config file). Works fully offline, no cloud speech APIs. - Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or osascript, then restore my previous clipboard). - Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in .env) to fix punctuation and drop filler words, then paste. If no key is set, this mode just does a plain paste. - Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10 transcripts with copy buttons. - Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the audio after transcription. No accounts, no telemetry. - Out of scope: per-app presets and custom vocabulary tuning. One good general mode. - README: mic + accessibility permissions to grant, how to download the whisper model, and a note that the first run will trigger macOS permission prompts. ## Required capabilities - macOS automation/accessibility permissions - speech API or local Whisper - optional LLM API - hotkey library ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# Superwhisper product brief ## Problem A push-to-talk recorder that transcribes and pastes text into the active app is very buildable for one person, especially on macOS. ## Product outcome Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - macOS automation/accessibility permissions - speech API or local Whisper - optional LLM API - hotkey library ## Explicit non-goals for v1 - beautiful native UX - presets - vocabulary/profile tuning - app-wide polish - support ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements: - A global hotkey (default: hold right Option) records my mic while held, stops on release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to ship. - Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp (small.en by default, model path in a config file). Works fully offline, no cloud speech APIs. - Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or osascript, then restore my previous clipboard). - Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in .env) to fix punctuation and drop filler words, then paste. If no key is set, this mode just does a plain paste. - Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10 transcripts with copy buttons. - Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the audio after transcription. No accounts, no telemetry. - Out of scope: per-app presets and custom vocabulary tuning. One good general mode. - README: mic + accessibility permissions to grant, how to download the whisper model, and a note that the first run will trigger macOS permission prompts. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted Superwhisper capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent · this prompt is generated from the build plan · improve it via PR
They pay for the low-friction menu-bar experience and dictation profiles.
xbeautiful native UX
xpresets
xvocabulary/profile tuning
xapp-wide polish
xsupport
Don't feel like building it? These folks already made it free.
all 8 free alternatives to Superwhisper →· no votes, no pay-to-list · just what's real
Superwhisper pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0 | $0 | Basic dictation; no local models or custom modes; current numeric cloud-usage cap is not published. |
| pro monthly | $8.49 | — | Unlimited cloud and local models; license covers personal Mac, Windows, iPhone and iPad devices. |
| pro annual | — | $7.08 | Unlimited cloud and local models; same personal-device license. |
| pro lifetime | custom | — | One-time $249.99 purchase; unlimited cloud and local models under the lifetime license terms. |
free tierBasic dictation; no local models or custom modes; the current numeric cloud-use allowance is not published.
billingmonthly, annual and one-time lifetime purchase; direct-web prices
hidden costsApp Store pricing is higher because of Apple fees, but the current exact App Store amounts were not published in the official Pro documentation.
verified 2026-08-11 · source ↗
Is Superwhisper free?
A free version covers basic dictation; Pro adds unlimited cloud and local models. Paid is Pro at $8.49/mo (checked 2026-08-07).
Vibecode Superwhisper
Yes. A competent AI coding agent (Claude Code, Codex, Cursor) can build a usable personal Superwhisper replacement in one session with the prompt on this page. It runs on your own machine or server with no subscription.
How much does Superwhisper cost?
Superwhisper costs about $8.49/month (Pro, checked 2026-08-07), which is $101.88 per year. That's what you save by replacing it with one prompt.
What do I lose by replacing Superwhisper?
Honestly: beautiful native UX; presets; vocabulary/profile tuning; app-wide polish; support. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to Superwhisper?
Yes: Voicebox (Voice cloning studio attached to a global dictation hotkey; overkill, free overkill.) Handy (Press a key, talk and get plain text in the app you were already using. No cleanup magic, no invoice.) FluidVoice (Fast Mac dictation with local cleanup; even the fancy part stays on the machine.) All 8 curated free alternatives are at vibecodeit.com/superwhisper/alternatives. The prompt is for when you want it exactly your way.