Vibecode ScanToExcel
track this build5 steps, step by step0%A vision model will read a clean printed table on the first try, and that demo is genuinely an afternoon of work. Consistency is the part that is not. Real documents arrive with merged cells, multi-line rows, columns that shift between pages and numbers that must survive as numbers, and a one-shot prompt handles each of those differently every time you run it. ScanToExcel puts a purpose-built extraction pipeline between the model and the spreadsheet precisely because the model alone is not reproducible. You can copy the easy half of this product in a sitting and spend months on the half that makes it trustworthy.
You are building a lean indie version of ScanToExcel. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # ScanToExcel indie build ## Goal Build the smallest trustworthy replacement for the core ScanToExcel workflow for one developer or a tiny team. ## Scope Upload a photo or PDF page, send it to a vision model asking for the table as structured rows, and write the result to an .xlsx file. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - a purpose-built extraction pipeline rather than raw model output - the same document producing the same spreadsheet twice - structure held across pages: merged cells, multi-line rows, shifting columns - reliable handwriting recognition - a phone app that captures and converts without a laptop If those capabilities are essential, use img2table instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a document-to-spreadsheet converter inspired by ScanToExcel. Use exactly this stack: Next.js 15 + TypeScript. Primary job: the user uploads a photo or a PDF page containing a table, and gets back a downloadable .xlsx whose cells match the document. Start from an empty folder and create the complete working project. Send the image to one vision-capable model and ask for the table as structured rows and columns, not as prose. Write the result with a spreadsheet library so numbers arrive as numbers and dates as dates, not as text. Show the extracted table in the browser for review and let the user fix cells before downloading - never hand over a file the user has not seen. Handle multi-page PDFs by processing pages in order and dropping a repeated header row when the same header appears on a later page. Single-user and private by default; delete uploads once the download is produced, and say so in the UI. Put every secret in .env and provide .env.example. Include clear empty, loading, validation, success and failure states, and show the model's confidence when it reports one. Accessible keyboard navigation, labels, focus states and sensible contrast. Deliberately exclude these paid-product advantages: a native mobile capture app, tuned handwriting recognition, batch processing of very long documents. State plainly in the README that handwriting and low-quality photos are where this build degrades. Write unit tests for the row-parsing logic and one end-to-end smoke test that converts a sample image to a valid .xlsx. Create a README with setup, architecture, per-page cost estimate and limitations. Run the tests and build before finishing, then fix what fails. ## Required capabilities - OpenAI or Anthropic API key with vision support, in .env - Node.js 22 - A spreadsheet writer such as SheetJS or ExcelJS ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of ScanToExcel. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # ScanToExcel indie build ## Goal Build the smallest trustworthy replacement for the core ScanToExcel workflow for one developer or a tiny team. ## Scope Upload a photo or PDF page, send it to a vision model asking for the table as structured rows, and write the result to an .xlsx file. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - a purpose-built extraction pipeline rather than raw model output - the same document producing the same spreadsheet twice - structure held across pages: merged cells, multi-line rows, shifting columns - reliable handwriting recognition - a phone app that captures and converts without a laptop If those capabilities are essential, use img2table instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build a document-to-spreadsheet converter inspired by ScanToExcel. Use exactly this stack: Next.js 15 + TypeScript. Primary job: the user uploads a photo or a PDF page containing a table, and gets back a downloadable .xlsx whose cells match the document. Start from an empty folder and create the complete working project. Send the image to one vision-capable model and ask for the table as structured rows and columns, not as prose. Write the result with a spreadsheet library so numbers arrive as numbers and dates as dates, not as text. Show the extracted table in the browser for review and let the user fix cells before downloading - never hand over a file the user has not seen. Handle multi-page PDFs by processing pages in order and dropping a repeated header row when the same header appears on a later page. Single-user and private by default; delete uploads once the download is produced, and say so in the UI. Put every secret in .env and provide .env.example. Include clear empty, loading, validation, success and failure states, and show the model's confidence when it reports one. Accessible keyboard navigation, labels, focus states and sensible contrast. Deliberately exclude these paid-product advantages: a native mobile capture app, tuned handwriting recognition, batch processing of very long documents. State plainly in the README that handwriting and low-quality photos are where this build degrades. Write unit tests for the row-parsing logic and one end-to-end smoke test that converts a sample image to a valid .xlsx. Create a README with setup, architecture, per-page cost estimate and limitations. Run the tests and build before finishing, then fix what fails. ## Required capabilities - OpenAI or Anthropic API key with vision support, in .env - Node.js 22 - A spreadsheet writer such as SheetJS or ExcelJS ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of ScanToExcel. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # ScanToExcel product brief ## Problem A vision model will read a clean printed table on the first try, and that demo is genuinely an afternoon of work. Consistency is the part that is not. Real documents arrive with merged cells, multi-line rows, columns that shift between pages and numbers that must survive as numbers, and a one-shot prompt handles each of those differently every time you run it. ScanToExcel puts a purpose-built extraction pipeline between the model and the spreadsheet precisely because the model alone is not reproducible. You can copy the easy half of this product in a sitting and spend months on the half that makes it trustworthy. ## Product outcome Upload a photo or PDF page, send it to a vision model asking for the table as structured rows, and write the result to an .xlsx file. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - OpenAI or Anthropic API key with vision support, in .env - Node.js 22 - A spreadsheet writer such as SheetJS or ExcelJS ## Explicit non-goals for v1 - a purpose-built extraction pipeline rather than raw model output - the same document producing the same spreadsheet twice - structure held across pages: merged cells, multi-line rows, shifting columns - reliable handwriting recognition - a phone app that captures and converts without a laptop ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build a document-to-spreadsheet converter inspired by ScanToExcel. Use exactly this stack: Next.js 15 + TypeScript. Primary job: the user uploads a photo or a PDF page containing a table, and gets back a downloadable .xlsx whose cells match the document. Start from an empty folder and create the complete working project. Send the image to one vision-capable model and ask for the table as structured rows and columns, not as prose. Write the result with a spreadsheet library so numbers arrive as numbers and dates as dates, not as text. Show the extracted table in the browser for review and let the user fix cells before downloading - never hand over a file the user has not seen. Handle multi-page PDFs by processing pages in order and dropping a repeated header row when the same header appears on a later page. Single-user and private by default; delete uploads once the download is produced, and say so in the UI. Put every secret in .env and provide .env.example. Include clear empty, loading, validation, success and failure states, and show the model's confidence when it reports one. Accessible keyboard navigation, labels, focus states and sensible contrast. Deliberately exclude these paid-product advantages: a native mobile capture app, tuned handwriting recognition, batch processing of very long documents. State plainly in the README that handwriting and low-quality photos are where this build degrades. Write unit tests for the row-parsing logic and one end-to-end smoke test that converts a sample image to a valid .xlsx. Create a README with setup, architecture, per-page cost estimate and limitations. Run the tests and build before finishing, then fix what fails. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted ScanToExcel capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# ScanToExcel indie build ## Goal Build the smallest trustworthy replacement for the core ScanToExcel workflow for one developer or a tiny team. ## Scope Upload a photo or PDF page, send it to a vision model asking for the table as structured rows, and write the result to an .xlsx file. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - a purpose-built extraction pipeline rather than raw model output - the same document producing the same spreadsheet twice - structure held across pages: merged cells, multi-line rows, shifting columns - reliable handwriting recognition - a phone app that captures and converts without a laptop If those capabilities are essential, use img2table instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build a document-to-spreadsheet converter inspired by ScanToExcel. Use exactly this stack: Next.js 15 + TypeScript. Primary job: the user uploads a photo or a PDF page containing a table, and gets back a downloadable .xlsx whose cells match the document. Start from an empty folder and create the complete working project. Send the image to one vision-capable model and ask for the table as structured rows and columns, not as prose. Write the result with a spreadsheet library so numbers arrive as numbers and dates as dates, not as text. Show the extracted table in the browser for review and let the user fix cells before downloading - never hand over a file the user has not seen. Handle multi-page PDFs by processing pages in order and dropping a repeated header row when the same header appears on a later page. Single-user and private by default; delete uploads once the download is produced, and say so in the UI. Put every secret in .env and provide .env.example. Include clear empty, loading, validation, success and failure states, and show the model's confidence when it reports one. Accessible keyboard navigation, labels, focus states and sensible contrast. Deliberately exclude these paid-product advantages: a native mobile capture app, tuned handwriting recognition, batch processing of very long documents. State plainly in the README that handwriting and low-quality photos are where this build degrades. Write unit tests for the row-parsing logic and one end-to-end smoke test that converts a sample image to a valid .xlsx. Create a README with setup, architecture, per-page cost estimate and limitations. Run the tests and build before finishing, then fix what fails. ## Required capabilities - OpenAI or Anthropic API key with vision support, in .env - Node.js 22 - A spreadsheet writer such as SheetJS or ExcelJS ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# ScanToExcel product brief ## Problem A vision model will read a clean printed table on the first try, and that demo is genuinely an afternoon of work. Consistency is the part that is not. Real documents arrive with merged cells, multi-line rows, columns that shift between pages and numbers that must survive as numbers, and a one-shot prompt handles each of those differently every time you run it. ScanToExcel puts a purpose-built extraction pipeline between the model and the spreadsheet precisely because the model alone is not reproducible. You can copy the easy half of this product in a sitting and spend months on the half that makes it trustworthy. ## Product outcome Upload a photo or PDF page, send it to a vision model asking for the table as structured rows, and write the result to an .xlsx file. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - OpenAI or Anthropic API key with vision support, in .env - Node.js 22 - A spreadsheet writer such as SheetJS or ExcelJS ## Explicit non-goals for v1 - a purpose-built extraction pipeline rather than raw model output - the same document producing the same spreadsheet twice - structure held across pages: merged cells, multi-line rows, shifting columns - reliable handwriting recognition - a phone app that captures and converts without a laptop ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build a document-to-spreadsheet converter inspired by ScanToExcel. Use exactly this stack: Next.js 15 + TypeScript. Primary job: the user uploads a photo or a PDF page containing a table, and gets back a downloadable .xlsx whose cells match the document. Start from an empty folder and create the complete working project. Send the image to one vision-capable model and ask for the table as structured rows and columns, not as prose. Write the result with a spreadsheet library so numbers arrive as numbers and dates as dates, not as text. Show the extracted table in the browser for review and let the user fix cells before downloading - never hand over a file the user has not seen. Handle multi-page PDFs by processing pages in order and dropping a repeated header row when the same header appears on a later page. Single-user and private by default; delete uploads once the download is produced, and say so in the UI. Put every secret in .env and provide .env.example. Include clear empty, loading, validation, success and failure states, and show the model's confidence when it reports one. Accessible keyboard navigation, labels, focus states and sensible contrast. Deliberately exclude these paid-product advantages: a native mobile capture app, tuned handwriting recognition, batch processing of very long documents. State plainly in the README that handwriting and low-quality photos are where this build degrades. Write unit tests for the row-parsing logic and one end-to-end smoke test that converts a sample image to a valid .xlsx. Create a README with setup, architecture, per-page cost estimate and limitations. Run the tests and build before finishing, then fix what fails. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted ScanToExcel capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent
Because the demo works and the hundredth document does not. Accountants feed it crooked phone photos of carbon-copy forms with merged headers, and the difference between a tool and a script is what happens on that page. Paying for output you can rely on without checking every cell is an easy trade for someone billing hourly.
xa purpose-built extraction pipeline rather than raw model output
xthe same document producing the same spreadsheet twice
xstructure held across pages: merged cells, multi-line rows, shifting columns
xreliable handwriting recognition
xa phone app that captures and converts without a laptop
ScanToExcel pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0 | $0 | 2 free pages on web plus 3 free scans on iOS; standard AI extraction; XLSX, CSV and JSON exports |
| page pack, 50 | custom | — | 50 pages; one-time $4.99; never expires |
| page pack, 200 | custom | — | 200 pages; one-time $14.99; never expires |
| pro | $39.99 | — | 1,000 pages/month; multi-page PDFs; batch uploads up to 20 files; priority processing |
| api | $59 | — | 1,000 API calls/month; REST API, webhooks, JSON and XLSX output; priority support/SLA |
free tier2 free pages on web plus 3 free scans on iOS; no sign-up required
billingmonthly subscriptions only for Pro/API; one-time non-expiring page packs also sold; cancel anytime
hidden costsAfter free pages are used, another scan requires Pro or a page pack. Paddle is merchant of record and calculates taxes at checkout. The $59 API tier is listed but marked Coming Soon.
verified 2026-08-13 · source ↗
Vibecode ScanToExcel
Kinda. The core of ScanToExcel is buildable in a weekend with the prompt on this page, but there are real gaps: a purpose-built extraction pipeline rather than raw model output, the same document producing the same spreadsheet twice. Read the honest list above before committing.
How much does ScanToExcel cost?
ScanToExcel costs about $39.99/month (Pro, checked 2026-08-03), which is $479.88 per year.
What do I lose by replacing ScanToExcel?
Honestly: a purpose-built extraction pipeline rather than raw model output; the same document producing the same spreadsheet twice; structure held across pages: merged cells, multi-line rows, shifting columns; reliable handwriting recognition; a phone app that captures and converts without a laptop. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to ScanToExcel?
Yes: img2table (Table identification and extraction from images and PDFs, no model API required.), Camelot (Extracts tables from text-based PDFs into DataFrames.), docling (Document parsing to structured formats, including table structure recovery.). Using prior art is also vibecoding; the prompt is for when you want it exactly your way.