Vibecode ekto
track this build5 steps, step by step0%The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market.
You are building a lean indie version of ekto. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # ekto indie build ## Goal Build the smallest trustworthy replacement for the core ekto workflow for one developer or a tiny team. ## Scope Streams mic audio over a WebSocket to a local Whisper plus translation plus TTS chain and plays the translated speech back with running transcript. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes - Clean sentence by sentence pacing and turn detection, which is most of the perceived quality - A mobile app, so no translating anything while standing up - Offline or low-bandwidth behavior tuned for actual travel - Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment If those capabilities are essential, use ekto instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a local real-time voice translation app. No accounts, no cloud services, no telemetry. Stack, non-negotiable: - Python 3.11 + FastAPI, served with uvicorn on port 8000. - One HTML page with vanilla JS, no framework, no build step. - Audio in: browser getUserMedia, 16kHz mono, streamed to the server over a WebSocket in 250ms PCM chunks. - Speech to text: faster-whisper (small model default, configurable via .env). - Segmentation: silero-vad or webrtcvad to detect end of utterance. Do not translate on fixed timers, translate on detected utterance boundaries. - Translation: argostranslate with locally installed language pairs. - Text to speech: piper, one voice per target language, downloaded on first run into ./models. Behavior: - User picks source and target language in a dropdown before starting. - Press Start, speak, and on each detected utterance the server returns: original text, translated text, and a WAV of the translated speech. The page appends both lines to a running transcript and plays the audio. - Show live latency per utterance in ms in the corner. Be honest, measure end of speech to audio ready. - Handle overlap: if a new utterance arrives while audio is playing, queue it, never drop it. - Long session hygiene: cap the in-memory transcript at 500 lines, reset the whisper buffer after every utterance, log RSS every 60 seconds. Out of scope, do not build: mobile app, user accounts, cloud sync, speaker diarization, a two-phone conversation mode. Deliverables: main.py, static/index.html, static/app.js, requirements.txt, .env.example (WHISPER_MODEL, DEVICE, COMPUTE_TYPE), scripts/download_models.py, and a README with exact run steps plus one paragraph on where this degrades in sessions over 20 minutes. Run it, speak a test sentence in English with Spanish as target, and paste the measured latency into the README. ## Required capabilities - Python 3.11 and a machine with at least 8GB RAM, GPU strongly preferred - Local model downloads: faster-whisper and a Piper voice per target language - A browser with mic permission, or an API key if you swap in a hosted translation model - Headphones, otherwise the TTS output feeds back into the mic ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of ekto. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # ekto indie build ## Goal Build the smallest trustworthy replacement for the core ekto workflow for one developer or a tiny team. ## Scope Streams mic audio over a WebSocket to a local Whisper plus translation plus TTS chain and plays the translated speech back with running transcript. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes - Clean sentence by sentence pacing and turn detection, which is most of the perceived quality - A mobile app, so no translating anything while standing up - Offline or low-bandwidth behavior tuned for actual travel - Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment If those capabilities are essential, use ekto instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a local real-time voice translation app. No accounts, no cloud services, no telemetry. Stack, non-negotiable: - Python 3.11 + FastAPI, served with uvicorn on port 8000. - One HTML page with vanilla JS, no framework, no build step. - Audio in: browser getUserMedia, 16kHz mono, streamed to the server over a WebSocket in 250ms PCM chunks. - Speech to text: faster-whisper (small model default, configurable via .env). - Segmentation: silero-vad or webrtcvad to detect end of utterance. Do not translate on fixed timers, translate on detected utterance boundaries. - Translation: argostranslate with locally installed language pairs. - Text to speech: piper, one voice per target language, downloaded on first run into ./models. Behavior: - User picks source and target language in a dropdown before starting. - Press Start, speak, and on each detected utterance the server returns: original text, translated text, and a WAV of the translated speech. The page appends both lines to a running transcript and plays the audio. - Show live latency per utterance in ms in the corner. Be honest, measure end of speech to audio ready. - Handle overlap: if a new utterance arrives while audio is playing, queue it, never drop it. - Long session hygiene: cap the in-memory transcript at 500 lines, reset the whisper buffer after every utterance, log RSS every 60 seconds. Out of scope, do not build: mobile app, user accounts, cloud sync, speaker diarization, a two-phone conversation mode. Deliverables: main.py, static/index.html, static/app.js, requirements.txt, .env.example (WHISPER_MODEL, DEVICE, COMPUTE_TYPE), scripts/download_models.py, and a README with exact run steps plus one paragraph on where this degrades in sessions over 20 minutes. Run it, speak a test sentence in English with Spanish as target, and paste the measured latency into the README. ## Required capabilities - Python 3.11 and a machine with at least 8GB RAM, GPU strongly preferred - Local model downloads: faster-whisper and a Piper voice per target language - A browser with mic permission, or an API key if you swap in a hosted translation model - Headphones, otherwise the TTS output feeds back into the mic ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of ekto. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # ekto product brief ## Problem The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market. ## Product outcome Streams mic audio over a WebSocket to a local Whisper plus translation plus TTS chain and plays the translated speech back with running transcript. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - Python 3.11 and a machine with at least 8GB RAM, GPU strongly preferred - Local model downloads: faster-whisper and a Piper voice per target language - A browser with mic permission, or an API key if you swap in a hosted translation model - Headphones, otherwise the TTS output feeds back into the mic ## Explicit non-goals for v1 - Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes - Clean sentence by sentence pacing and turn detection, which is most of the perceived quality - A mobile app, so no translating anything while standing up - Offline or low-bandwidth behavior tuned for actual travel - Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build a local real-time voice translation app. No accounts, no cloud services, no telemetry. Stack, non-negotiable: - Python 3.11 + FastAPI, served with uvicorn on port 8000. - One HTML page with vanilla JS, no framework, no build step. - Audio in: browser getUserMedia, 16kHz mono, streamed to the server over a WebSocket in 250ms PCM chunks. - Speech to text: faster-whisper (small model default, configurable via .env). - Segmentation: silero-vad or webrtcvad to detect end of utterance. Do not translate on fixed timers, translate on detected utterance boundaries. - Translation: argostranslate with locally installed language pairs. - Text to speech: piper, one voice per target language, downloaded on first run into ./models. Behavior: - User picks source and target language in a dropdown before starting. - Press Start, speak, and on each detected utterance the server returns: original text, translated text, and a WAV of the translated speech. The page appends both lines to a running transcript and plays the audio. - Show live latency per utterance in ms in the corner. Be honest, measure end of speech to audio ready. - Handle overlap: if a new utterance arrives while audio is playing, queue it, never drop it. - Long session hygiene: cap the in-memory transcript at 500 lines, reset the whisper buffer after every utterance, log RSS every 60 seconds. Out of scope, do not build: mobile app, user accounts, cloud sync, speaker diarization, a two-phone conversation mode. Deliverables: main.py, static/index.html, static/app.js, requirements.txt, .env.example (WHISPER_MODEL, DEVICE, COMPUTE_TYPE), scripts/download_models.py, and a README with exact run steps plus one paragraph on where this degrades in sessions over 20 minutes. Run it, speak a test sentence in English with Spanish as target, and paste the measured latency into the README. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted ekto capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# ekto indie build ## Goal Build the smallest trustworthy replacement for the core ekto workflow for one developer or a tiny team. ## Scope Streams mic audio over a WebSocket to a local Whisper plus translation plus TTS chain and plays the translated speech back with running transcript. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes - Clean sentence by sentence pacing and turn detection, which is most of the perceived quality - A mobile app, so no translating anything while standing up - Offline or low-bandwidth behavior tuned for actual travel - Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment If those capabilities are essential, use ekto instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build a local real-time voice translation app. No accounts, no cloud services, no telemetry. Stack, non-negotiable: - Python 3.11 + FastAPI, served with uvicorn on port 8000. - One HTML page with vanilla JS, no framework, no build step. - Audio in: browser getUserMedia, 16kHz mono, streamed to the server over a WebSocket in 250ms PCM chunks. - Speech to text: faster-whisper (small model default, configurable via .env). - Segmentation: silero-vad or webrtcvad to detect end of utterance. Do not translate on fixed timers, translate on detected utterance boundaries. - Translation: argostranslate with locally installed language pairs. - Text to speech: piper, one voice per target language, downloaded on first run into ./models. Behavior: - User picks source and target language in a dropdown before starting. - Press Start, speak, and on each detected utterance the server returns: original text, translated text, and a WAV of the translated speech. The page appends both lines to a running transcript and plays the audio. - Show live latency per utterance in ms in the corner. Be honest, measure end of speech to audio ready. - Handle overlap: if a new utterance arrives while audio is playing, queue it, never drop it. - Long session hygiene: cap the in-memory transcript at 500 lines, reset the whisper buffer after every utterance, log RSS every 60 seconds. Out of scope, do not build: mobile app, user accounts, cloud sync, speaker diarization, a two-phone conversation mode. Deliverables: main.py, static/index.html, static/app.js, requirements.txt, .env.example (WHISPER_MODEL, DEVICE, COMPUTE_TYPE), scripts/download_models.py, and a README with exact run steps plus one paragraph on where this degrades in sessions over 20 minutes. Run it, speak a test sentence in English with Spanish as target, and paste the measured latency into the README. ## Required capabilities - Python 3.11 and a machine with at least 8GB RAM, GPU strongly preferred - Local model downloads: faster-whisper and a Piper voice per target language - A browser with mic permission, or an API key if you swap in a hosted translation model - Headphones, otherwise the TTS output feeds back into the mic ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# ekto product brief ## Problem The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market. ## Product outcome Streams mic audio over a WebSocket to a local Whisper plus translation plus TTS chain and plays the translated speech back with running transcript. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - Python 3.11 and a machine with at least 8GB RAM, GPU strongly preferred - Local model downloads: faster-whisper and a Piper voice per target language - A browser with mic permission, or an API key if you swap in a hosted translation model - Headphones, otherwise the TTS output feeds back into the mic ## Explicit non-goals for v1 - Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes - Clean sentence by sentence pacing and turn detection, which is most of the perceived quality - A mobile app, so no translating anything while standing up - Offline or low-bandwidth behavior tuned for actual travel - Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build a local real-time voice translation app. No accounts, no cloud services, no telemetry. Stack, non-negotiable: - Python 3.11 + FastAPI, served with uvicorn on port 8000. - One HTML page with vanilla JS, no framework, no build step. - Audio in: browser getUserMedia, 16kHz mono, streamed to the server over a WebSocket in 250ms PCM chunks. - Speech to text: faster-whisper (small model default, configurable via .env). - Segmentation: silero-vad or webrtcvad to detect end of utterance. Do not translate on fixed timers, translate on detected utterance boundaries. - Translation: argostranslate with locally installed language pairs. - Text to speech: piper, one voice per target language, downloaded on first run into ./models. Behavior: - User picks source and target language in a dropdown before starting. - Press Start, speak, and on each detected utterance the server returns: original text, translated text, and a WAV of the translated speech. The page appends both lines to a running transcript and plays the audio. - Show live latency per utterance in ms in the corner. Be honest, measure end of speech to audio ready. - Handle overlap: if a new utterance arrives while audio is playing, queue it, never drop it. - Long session hygiene: cap the in-memory transcript at 500 lines, reset the whisper buffer after every utterance, log RSS every 60 seconds. Out of scope, do not build: mobile app, user accounts, cloud sync, speaker diarization, a two-phone conversation mode. Deliverables: main.py, static/index.html, static/app.js, requirements.txt, .env.example (WHISPER_MODEL, DEVICE, COMPUTE_TYPE), scripts/download_models.py, and a README with exact run steps plus one paragraph on where this degrades in sessions over 20 minutes. Run it, speak a test sentence in English with Spanish as target, and paste the measured latency into the README. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted ekto capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent · this prompt is generated from the build plan · improve it via PR
Because voice translation is judged entirely on the seconds between someone finishing a sentence and you hearing it, and on whether it still works on minute 40. A local build nails the demo and then frays: barge-in, background noise, two people talking over each other, the phone locking. Paying gets you a phone in your pocket that handles those cases without you adding VAD thresholds mid-conversation.
xLong session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes
xClean sentence by sentence pacing and turn detection, which is most of the perceived quality
xA mobile app, so no translating anything while standing up
xOffline or low-bandwidth behavior tuned for actual travel
xLatency budgets someone else already fought for: streaming partial results instead of waiting for a full segment
Nothing worth pointing at. That's why the prompt exists.
Vibecode ekto
Kinda. The core of ekto is buildable in a weekend with the prompt on this page, but there are real gaps: Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes, Clean sentence by sentence pacing and turn detection, which is most of the perceived quality. Read the honest list above before committing.
How much does ekto cost?
ekto costs about $29.99/month (Monthly Unlimited PRO, checked 2026-08-18), which is $359.88 per year.
What do I lose by replacing ekto?
Honestly: Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes; Clean sentence by sentence pacing and turn detection, which is most of the perceived quality; A mobile app, so no translating anything while standing up; Offline or low-bandwidth behavior tuned for actual travel; Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to ekto?
No mature open-source alternative worth pointing at, which is exactly why the one-shot prompt on this page exists.