Implementation overview
How catalog data is handled
This is an implementation overview, not a privacy policy. It describes the current catalog workflow and its technical boundaries.
Catalog sources and workspace records
Creating a catalog job submits a source file and creates source and job records. Product records are created during processing; review, mapping, and export records are created only when those actions occur. Source artifacts are stored in a private catalog-artifacts bucket and access is scoped through organization membership policies.
This page describes the product implementation. It does not provide legal terms, provider policies, or data-residency commitments.
OCR processing
For the Azure OCR path, document bytes are sent as base64 data to the configured Azure Mistral Document AI endpoint so document content can be processed.
This implementation overview does not make claims about provider retention, training, or regional processing policies.
Source-artifact retention
Catalog source artifacts and processing checkpoints are configured with a default retention window of 30 days. After an eligible completed or cancelled workflow, the catalog worker can purge those source artifacts and mark the source as purged.
This does not describe retention of account, review, or export metadata. Those records support the catalog workflow and have separate lifecycle behavior.
Review and export history
Review decisions and catalog exports keep workflow history beside the records they concern. Approved-only export prevents unapproved catalog records from being included in a CSV or JSON download.
For questions about a specific workspace or source file, use the product support channel available to your team.
Optional product analytics
Analytics is disabled until a visitor allows it. After consent, toSchema records pathname-only page views, a coarse acquisition channel, calls to action, and successful signup, upload, review, and export stages.
The event fields can include user, organization, source, job, product, mapping, and export identifiers plus fixed status, format, and count fields. They exclude filenames, document content, extracted values, evidence, full URLs, referrers, and campaign values. PostHog receives the product events, while Vercel Analytics receives pathname-only page events.