AI data cleansing for mid-market teams

Clean data, approved by your people

Connect a warehouse or drop a CSV. Data Stew profiles every column, finds the issues, and proposes fixes with rationale and evidence — nothing touches your data until a person approves it.

Free tier includes 10 AI fix suggestions and a shareable quality report. No card required.

Review queuecustomers.csv · 48,112 rows
Score 74 → 964 of 10 suggestions
ColumnCurrentProposedConfidence
stateN.S.W.NSW98%
segment∅ missingWholesale93%
phone412345678+61 412 345 67896%
abn51 824 753 5— abstainedneeds you

Every suggestion carries its model, prompt version, and evidence rows — an audit trail your security team can export.

Anthropic under BAA · Mistral under zero-data-retention

Provider allowlist enforced in code, not policy PDFs

Your data stays in your warehouse — we hold metadata

What it does

The whole loop, not another dashboard of alerts

Detection tools tell you what’s broken and stop. Data Stew owns the fix: suggestion, approval, writeback, and the score that proves it worked.

01

Profiles everything. Flags what matters.

Null rates, format distributions, near-duplicate categories, cross-column dependencies — computed as push-down SQL in your warehouse, or DuckDB for files. Rules are generated from what your data actually looks like, and every dataset gets an explainable quality score.

  • Completeness99.2%
  • Format validity96.4%
  • Uniqueness100%
  • Cross-field consistency94.1%

02

Proposes fixes with reasons, not magic.

AI looks at similar records the way a person would, then proposes a value with its confidence, rationale, and the evidence rows it used. When the evidence is thin, it abstains — a wrong guess costs more trust than a blank.

SydnySydney 97%

“14 of 15 records at this postcode read Sydney; the outlier matches no known suburb.”

03

Your stewards stay in charge.

A review grid built for throughput: approve, edit, or reject; bulk-accept everything above a confidence bar. Systematic fixes are approved once as a pattern — one decision can correct thousands of rows.

Pattern: canonicalize state namesapproved

Rows corrected by this decision312

04

Every change is on the record.

Fixes land as a suggestion table and a script your team runs — or export a cleaned file. Who approved what, which model proposed it, under which prompt version: an append-only audit log built for your auditors.

  • apply.completed · 312 cells · approved by M. Chen
  • suggestion.generated · claude-sonnet · canonicalize@v1
  • profile.completed · customers.csv · score 74

How it works

Minutes to set up. First results today.

1

Drop a file or connect

Upload a CSV, or connect Postgres, MySQL, SQL Server, Snowflake, or BigQuery with a read-only role — the least-privilege setup is generated for you.

2

Get your quality report

About a minute later: a score, a per-column issue breakdown, and sample AI fixes with their reasoning. Shareable with anyone, no account needed.

3

Approve and apply

Work the review queue, bulk-accept the confident ones, and export the cleaned file or the fix script for your warehouse.

~60 sec

from CSV upload to quality report

10,000+

cells one approved pattern can fix

2

AI providers allowed to see cell data

Cell-level data goes only to Anthropic (BAA) or Mistral (zero data retention) — a provider allowlist the code refuses to violate.

Why teams switch

Built against the ways this usually fails

vs Legacy DQ suites

Six-month implementations, six-figure contracts, a services team on site.

Connect to first fix suggestions in under an hour, from ~$500 a month. No implementation project.

vs Detection-only tools

Dev-centric alerting: you learn what’s broken, and the tickets pile up.

We own the fix — suggestion, steward approval, writeback — operated by the business, not the backlog.

vs Black-box ML mastering

A model changes your records and asks you to trust it.

Every fix carries its rationale, evidence, and approval trail. AI proposes, your people dispose.

Pricing

Simple pricing, metered on records

Start free with one file and ten suggested fixes. Paid plans meter on records suggested, with credit top-ups when a big cleanup runs long.

Free

$0 / month

See what’s wrong with one file, with the reasoning behind every fix.

  • Profile one CSV and score it
  • 10 sample fixes with rationale and evidence
  • Approve them and export the cleaned file
  • Shareable report (watermarked)
Start free

Starter

Most popular

$500 / month

One source, a steward and a reviewer, the full approve-and-apply loop.

  • 2 connected sources, 2 seats
  • 25,000 suggested records a month
  • Approve, edit, reject — then export a cleaned file
  • Full audit trail, kept 12 months
Get started

Team

$1500 / month

Several sources, a steward per domain, and the dashboards to run it.

  • 5 connected sources, 10 seats
  • 150,000 suggested records a month
  • Domains, stewards, and per-domain scores
  • Audit trail kept 24 months
Get started

Enterprise

Let's talk

EU residency, BYOK, SSO, and a contract your security team has read.

  • Unlimited sources and seats
  • Per-workspace provider pinning, EU residency, BYOK
  • SSO/SAML, custom DPA, security review support
Talk to us

Your data isn’t getting cleaner on its own

Upload one messy file and see the score, the issues, and ten proposed fixes — with the reasoning — in about a minute.