Clean before AI
Changelog
Every important Clean before AI release: what we improved, how the measurements changed and what still doesn't work well. We don't hide weaknesses.
Current state
98.2 %
Regular documents
95.6 %
Hard documents
98.2 %
Real public documents
92.9 %
Blind test
Rules + local model, 433 test documents in 17 languages, 0 false hits on texts without personal data. Measured 2026-10-03.
- v1.33 October 2026 · web tool + Chrome extension
Our own local AI model, known names and a confidence light
- Our own AI model (~53 MB, 17 languages) finds names, nicknames, addresses and companies the rules don't know. Bundled in the extension and works offline; in the web tool it downloads once when switched on and stays in your browser. Replaces the previous optional AI layer.
- Known names: colleagues, clients and projects are always hidden — in any form, in capitals and without accents. CSV import; organisations can distribute the list through a Chrome policy.
- Confidence light: doubtful words appear in a “Check” row — one click keeps them in the text.
- Detection: chats (speakers, nicknames, Cyrillic), declined addresses, small businesses without a legal form, form field labels, gaps found by the blind test.
Measurements
Regular documents 96.7 % → 98.2 % Hard documents 89.8 % → 95.6 % Real public forms and contracts (new) 98.2 % Blind test — texts the engine had never seen (new) 92.9 % False hits on texts without personal data 0 → 0 Known limitations
- On unknown texts (blind test) roughly one item in 14 stays unmarked. Review the highlighted parts before sending.
- Informal chats are the weakest area: nicknames, lowercase names, typos.
- Some languages (English, French, Slovak) have only a few hard test documents, so their hard-document figure is not yet reliable.
- v1.2.12 October 2026 · Chrome extension
Faster pasting in Gemini
- Gemini: long text (37,000 characters) pastes in 0.2 s instead of 6.2 s; the cursor stays after the text.
- Companies written as “D.O.O” without the final dot and “D.0.0.” from scans.
- Terms of use and a “Support the project” link in the panel.
- v1.230 September 2026 · web tool + Chrome extension
17 languages and real documents
- New languages: Polish, Dutch, Romanian, Czech, Slovak, Hungarian; English for the UK and US. Language is detected per paragraph, labels follow the language of the text.
- Extension interface in 14 languages.
- Real documents (public contracts and forms): domestic account numbers, IBANs across line breaks, health insurance numbers, phones, forms, layout-aware PDF reading; fewer false hits.
- “Also hide amounts” switch; the side panel opens by itself when cleaning.
Measurements
Test documents 136 → 263 Regular documents 96.5 % → 96.7 % Hard documents 89.1 % → 89.8 % False hits on texts without personal data 1 → 0 - v1.128 September 2026 · web tool + Chrome extension
First public measurements
- “How well it works” section: share of personal data found on test documents the engine had not seen, and false hits — re-checked after every change.
- Spanish and Portuguese, passwords and keys, server names, references, tables.
Measurements
Regular documents (136) 96.5 % Hard documents (28) 89.1 % - v1.027 September 2026 · web tool + Chrome extension
First version
- Local cleaning of personal data in the browser: declined names, consistent labels, real data restored in the AI's answer.
- IDs with checksums (national IDs, tax numbers, IBAN …) for Slovenia, the Balkans, Germany, Austria, Italy and France.
- PDF and image export with redacted areas.
- Chrome extension: ChatGPT, Claude, Gemini, Copilot, Perplexity, DeepSeek, Mistral, Grok and more; real data only in the side panel.