Vibecode VidNotes
track this build5 steps, step by step0%The narrow DIY loop is achievable in a weekend: import a file or YouTube URL, transcribe it, generate notes, and save the result locally. VidNotes as a product is broader than that prototype. It maintains ingestion across changing video sites, runs long jobs reliably, ships iOS, Android, web, and a Chrome extension, keeps account state consistent, and exposes the workflow to agents through an API, CLI, and MCP server. Rebuilding all of that is a multi-day integration project with an ongoing operations burden.
You are building a lean indie version of VidNotes. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # VidNotes indie build ## Goal Build the smallest trustworthy replacement for the core VidNotes workflow for one developer or a tiny team. ## Scope Import a video file or YouTube URL, extract and transcribe its audio, generate structured study notes, search the timestamped transcript, and export the result. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - native iOS and Android apps, a Chrome extension, and cross-device Premium state - managed extraction for TikTok, Instagram, Vimeo, and changing source sites - background processing, queue reliability, and long-video handling - 30+ language polish and prebuilt flashcards, quizzes, and chat - hosted web, API, CLI, and MCP access for agents If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local video transcription and study-notes app to replace VidNotes. Requirements: - A Python 3.12 app using FastAPI, Jinja templates, HTMX, and SQLite FTS5, served only on localhost:8787. - The home screen accepts one YouTube URL or an MP4, MOV, M4A, or MP3 upload; use yt-dlp for the URL and ffmpeg to extract mono audio. - Use the OpenAI Python SDK for timestamped speech-to-text and the Responses API; keep the API key in .env and split large audio into chunks with preserved offsets. - Show queued, extracting, transcribing, generating, complete, and failed states. The result page has a player plus a searchable transcript with clickable timestamps. - Generate a 5-bullet summary, key points, 10 flashcards, and a 5-question quiz from the completed transcript, with a button to regenerate each section. - Store metadata and generated text in SQLite, media under ./data/projects/<uuid>/, and export the transcript and notes as Markdown, TXT, SRT, and JSON. - No accounts and no telemetry. Everything stays local except yt-dlp downloads and OpenAI API calls, and each project keeps its original source URL. - Out of scope: native apps, hosted web, cross-device sync, the Chrome extension, Instagram or TikTok extraction, and a public API, CLI, or MCP server. - Include a README with setup, ffmpeg and yt-dlp installation, .env.example, data location, deletion steps, and a warning that source-site changes can break imports. ## Required capabilities - Python 3.12 - ffmpeg - yt-dlp - OpenAI API key - local disk ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of VidNotes. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # VidNotes indie build ## Goal Build the smallest trustworthy replacement for the core VidNotes workflow for one developer or a tiny team. ## Scope Import a video file or YouTube URL, extract and transcribe its audio, generate structured study notes, search the timestamped transcript, and export the result. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - native iOS and Android apps, a Chrome extension, and cross-device Premium state - managed extraction for TikTok, Instagram, Vimeo, and changing source sites - background processing, queue reliability, and long-video handling - 30+ language polish and prebuilt flashcards, quizzes, and chat - hosted web, API, CLI, and MCP access for agents If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local video transcription and study-notes app to replace VidNotes. Requirements: - A Python 3.12 app using FastAPI, Jinja templates, HTMX, and SQLite FTS5, served only on localhost:8787. - The home screen accepts one YouTube URL or an MP4, MOV, M4A, or MP3 upload; use yt-dlp for the URL and ffmpeg to extract mono audio. - Use the OpenAI Python SDK for timestamped speech-to-text and the Responses API; keep the API key in .env and split large audio into chunks with preserved offsets. - Show queued, extracting, transcribing, generating, complete, and failed states. The result page has a player plus a searchable transcript with clickable timestamps. - Generate a 5-bullet summary, key points, 10 flashcards, and a 5-question quiz from the completed transcript, with a button to regenerate each section. - Store metadata and generated text in SQLite, media under ./data/projects/<uuid>/, and export the transcript and notes as Markdown, TXT, SRT, and JSON. - No accounts and no telemetry. Everything stays local except yt-dlp downloads and OpenAI API calls, and each project keeps its original source URL. - Out of scope: native apps, hosted web, cross-device sync, the Chrome extension, Instagram or TikTok extraction, and a public API, CLI, or MCP server. - Include a README with setup, ffmpeg and yt-dlp installation, .env.example, data location, deletion steps, and a warning that source-site changes can break imports. ## Required capabilities - Python 3.12 - ffmpeg - yt-dlp - OpenAI API key - local disk ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of VidNotes. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # VidNotes product brief ## Problem The narrow DIY loop is achievable in a weekend: import a file or YouTube URL, transcribe it, generate notes, and save the result locally. VidNotes as a product is broader than that prototype. It maintains ingestion across changing video sites, runs long jobs reliably, ships iOS, Android, web, and a Chrome extension, keeps account state consistent, and exposes the workflow to agents through an API, CLI, and MCP server. Rebuilding all of that is a multi-day integration project with an ongoing operations burden. ## Product outcome Import a video file or YouTube URL, extract and transcribe its audio, generate structured study notes, search the timestamped transcript, and export the result. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - Python 3.12 - ffmpeg - yt-dlp - OpenAI API key - local disk ## Explicit non-goals for v1 - native iOS and Android apps, a Chrome extension, and cross-device Premium state - managed extraction for TikTok, Instagram, Vimeo, and changing source sites - background processing, queue reliability, and long-video handling - 30+ language polish and prebuilt flashcards, quizzes, and chat - hosted web, API, CLI, and MCP access for agents ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build me a local video transcription and study-notes app to replace VidNotes. Requirements: - A Python 3.12 app using FastAPI, Jinja templates, HTMX, and SQLite FTS5, served only on localhost:8787. - The home screen accepts one YouTube URL or an MP4, MOV, M4A, or MP3 upload; use yt-dlp for the URL and ffmpeg to extract mono audio. - Use the OpenAI Python SDK for timestamped speech-to-text and the Responses API; keep the API key in .env and split large audio into chunks with preserved offsets. - Show queued, extracting, transcribing, generating, complete, and failed states. The result page has a player plus a searchable transcript with clickable timestamps. - Generate a 5-bullet summary, key points, 10 flashcards, and a 5-question quiz from the completed transcript, with a button to regenerate each section. - Store metadata and generated text in SQLite, media under ./data/projects/<uuid>/, and export the transcript and notes as Markdown, TXT, SRT, and JSON. - No accounts and no telemetry. Everything stays local except yt-dlp downloads and OpenAI API calls, and each project keeps its original source URL. - Out of scope: native apps, hosted web, cross-device sync, the Chrome extension, Instagram or TikTok extraction, and a public API, CLI, or MCP server. - Include a README with setup, ffmpeg and yt-dlp installation, .env.example, data location, deletion steps, and a warning that source-site changes can break imports. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted VidNotes capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# VidNotes indie build ## Goal Build the smallest trustworthy replacement for the core VidNotes workflow for one developer or a tiny team. ## Scope Import a video file or YouTube URL, extract and transcribe its audio, generate structured study notes, search the timestamped transcript, and export the result. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - native iOS and Android apps, a Chrome extension, and cross-device Premium state - managed extraction for TikTok, Instagram, Vimeo, and changing source sites - background processing, queue reliability, and long-video handling - 30+ language polish and prebuilt flashcards, quizzes, and chat - hosted web, API, CLI, and MCP access for agents If those capabilities are essential, use whisper.cpp instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build me a local video transcription and study-notes app to replace VidNotes. Requirements: - A Python 3.12 app using FastAPI, Jinja templates, HTMX, and SQLite FTS5, served only on localhost:8787. - The home screen accepts one YouTube URL or an MP4, MOV, M4A, or MP3 upload; use yt-dlp for the URL and ffmpeg to extract mono audio. - Use the OpenAI Python SDK for timestamped speech-to-text and the Responses API; keep the API key in .env and split large audio into chunks with preserved offsets. - Show queued, extracting, transcribing, generating, complete, and failed states. The result page has a player plus a searchable transcript with clickable timestamps. - Generate a 5-bullet summary, key points, 10 flashcards, and a 5-question quiz from the completed transcript, with a button to regenerate each section. - Store metadata and generated text in SQLite, media under ./data/projects/<uuid>/, and export the transcript and notes as Markdown, TXT, SRT, and JSON. - No accounts and no telemetry. Everything stays local except yt-dlp downloads and OpenAI API calls, and each project keeps its original source URL. - Out of scope: native apps, hosted web, cross-device sync, the Chrome extension, Instagram or TikTok extraction, and a public API, CLI, or MCP server. - Include a README with setup, ffmpeg and yt-dlp installation, .env.example, data location, deletion steps, and a warning that source-site changes can break imports. ## Required capabilities - Python 3.12 - ffmpeg - yt-dlp - OpenAI API key - local disk ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# VidNotes product brief ## Problem The narrow DIY loop is achievable in a weekend: import a file or YouTube URL, transcribe it, generate notes, and save the result locally. VidNotes as a product is broader than that prototype. It maintains ingestion across changing video sites, runs long jobs reliably, ships iOS, Android, web, and a Chrome extension, keeps account state consistent, and exposes the workflow to agents through an API, CLI, and MCP server. Rebuilding all of that is a multi-day integration project with an ongoing operations burden. ## Product outcome Import a video file or YouTube URL, extract and transcribe its audio, generate structured study notes, search the timestamped transcript, and export the result. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - Python 3.12 - ffmpeg - yt-dlp - OpenAI API key - local disk ## Explicit non-goals for v1 - native iOS and Android apps, a Chrome extension, and cross-device Premium state - managed extraction for TikTok, Instagram, Vimeo, and changing source sites - background processing, queue reliability, and long-video handling - 30+ language polish and prebuilt flashcards, quizzes, and chat - hosted web, API, CLI, and MCP access for agents ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build me a local video transcription and study-notes app to replace VidNotes. Requirements: - A Python 3.12 app using FastAPI, Jinja templates, HTMX, and SQLite FTS5, served only on localhost:8787. - The home screen accepts one YouTube URL or an MP4, MOV, M4A, or MP3 upload; use yt-dlp for the URL and ffmpeg to extract mono audio. - Use the OpenAI Python SDK for timestamped speech-to-text and the Responses API; keep the API key in .env and split large audio into chunks with preserved offsets. - Show queued, extracting, transcribing, generating, complete, and failed states. The result page has a player plus a searchable transcript with clickable timestamps. - Generate a 5-bullet summary, key points, 10 flashcards, and a 5-question quiz from the completed transcript, with a button to regenerate each section. - Store metadata and generated text in SQLite, media under ./data/projects/<uuid>/, and export the transcript and notes as Markdown, TXT, SRT, and JSON. - No accounts and no telemetry. Everything stays local except yt-dlp downloads and OpenAI API calls, and each project keeps its original source URL. - Out of scope: native apps, hosted web, cross-device sync, the Chrome extension, Instagram or TikTok extraction, and a public API, CLI, or MCP server. - Include a README with setup, ffmpeg and yt-dlp installation, .env.example, data location, deletion steps, and a warning that source-site changes can break imports. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted VidNotes capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent
People pay for the maintained system around transcription: source adapters that keep working, long-job processing and retries, consistent results across native apps, web, and the extension, synced account state, and stable API, CLI, and MCP access for agents. The recurring value is avoiding integration breakage and operations work, not just getting speech-to-text.
xnative iOS and Android apps, a Chrome extension, and cross-device Premium state
xmanaged extraction for TikTok, Instagram, Vimeo, and changing source sites
xbackground processing, queue reliability, and long-video handling
x30+ language polish and prebuilt flashcards, quizzes, and chat
xhosted web, API, CLI, and MCP access for agents
Don't feel like building it? These folks already made it free.
no votes, no pay-to-list · just what's real
VidNotes pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| monthly | $9.99 | $4.17 | Unlimited transcription, summaries, flashcards, chat and action items; 30+ languages |
free tierno free tier; any trial allowance is not numerically disclosed on the public pricing section
billingmonthly or $49.99/year; no per-minute transcription charge advertised
verified 2026-08-14 · source ↗
Vibecode VidNotes
Kinda. The core of VidNotes is buildable in a weekend with the prompt on this page, but there are real gaps: native iOS and Android apps, a Chrome extension, and cross-device Premium state, managed extraction for TikTok, Instagram, Vimeo, and changing source sites. Read the honest list above before committing.
How much does VidNotes cost?
VidNotes costs about $9.99/month (Monthly, checked 2026-07-31), which is $119.88 per year.
What do I lose by replacing VidNotes?
Honestly: native iOS and Android apps, a Chrome extension, and cross-device Premium state; managed extraction for TikTok, Instagram, Vimeo, and changing source sites; background processing, queue reliability, and long-video handling; 30+ language polish and prebuilt flashcards, quizzes, and chat; hosted web, API, CLI, and MCP access for agents. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to VidNotes?
Yes: StoryToolkitAI (Local transcription and semantic footage search from an app that still admits it is rough.) NotebookLM (Drop in a YouTube video and get searchable, source-grounded study material; Google keeps the notebook.) The prompt is for when you want it exactly your way.