Vibecode Cluely
track this build5 steps, step by step0%The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters.
You are building a lean indie version of Cluely. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # Cluely indie build ## Goal Build the smallest trustworthy replacement for the core Cluely workflow for one developer or a tiny team. ## Scope A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools - Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking - System audio capture that just works without you installing and routing a virtual audio device - Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine - Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything' If those capabilities are essential, use Cluely instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing. ## Required capabilities - An OpenAI API key (or any multimodal chat endpoint) in .env - macOS or Windows desktop, plus screen recording and microphone permissions - A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows - Node 20 and willingness to sign or self-trust an unsigned Electron build ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of Cluely. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # Cluely indie build ## Goal Build the smallest trustworthy replacement for the core Cluely workflow for one developer or a tiny team. ## Scope A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools - Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking - System audio capture that just works without you installing and routing a virtual audio device - Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine - Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything' If those capabilities are essential, use Cluely instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing. ## Required capabilities - An OpenAI API key (or any multimodal chat endpoint) in .env - macOS or Windows desktop, plus screen recording and microphone permissions - A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows - Node 20 and willingness to sign or self-trust an unsigned Electron build ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of Cluely. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # Cluely product brief ## Problem The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters. ## Product outcome A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - An OpenAI API key (or any multimodal chat endpoint) in .env - macOS or Windows desktop, plus screen recording and microphone permissions - A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows - Node 20 and willingness to sign or self-trust an unsigned Electron build ## Explicit non-goals for v1 - Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools - Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking - System audio capture that just works without you installing and routing a virtual audio device - Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine - Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything' ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted Cluely capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# Cluely indie build ## Goal Build the smallest trustworthy replacement for the core Cluely workflow for one developer or a tiny team. ## Scope A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools - Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking - System audio capture that just works without you installing and routing a virtual audio device - Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine - Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything' If those capabilities are essential, use Cluely instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing. ## Required capabilities - An OpenAI API key (or any multimodal chat endpoint) in .env - macOS or Windows desktop, plus screen recording and microphone permissions - A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows - Node 20 and willingness to sign or self-trust an unsigned Electron build ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# Cluely product brief ## Problem The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters. ## Product outcome A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - An OpenAI API key (or any multimodal chat endpoint) in .env - macOS or Windows desktop, plus screen recording and microphone permissions - A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows - Node 20 and willingness to sign or self-trust an unsigned Electron build ## Explicit non-goals for v1 - Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools - Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking - System audio capture that just works without you installing and routing a virtual audio device - Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine - Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything' ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted Cluely capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent · this prompt is generated from the build plan · improve it via PR
Because the hard part is not the LLM call, it is the twenty small platform details that make an overlay silent, invisible and fast under pressure, and because people who reach for this tool are, by definition, not in the mood to debug a virtual audio driver ten minutes before an interview. Paying converts a fragile personal hack into something that mostly behaves on a laptop you did not configure yourself.
xScreen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools
xSub-second latency and the prompt tuning that makes answers short enough to read while someone is talking
xSystem audio capture that just works without you installing and routing a virtual audio device
xMobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine
xSomebody else's legal and PR exposure for a product whose pitch is 'cheat on everything'
Nothing worth pointing at. That's why the prompt exists.
Vibecode Cluely
Kinda. The core of Cluely is buildable in a weekend with the prompt on this page, but there are real gaps: Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools, Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking. Read the honest list above before committing.
How much does Cluely cost?
Cluely costs about $19.99/month (Pro, checked 2026-08-16), which is $239.88 per year.
What do I lose by replacing Cluely?
Honestly: Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools; Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking; System audio capture that just works without you installing and routing a virtual audio device; Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine; Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything'. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to Cluely?
No mature open-source alternative worth pointing at, which is exactly why the one-shot prompt on this page exists.