Clean before AI

Changelog

Every important Clean before AI release: what we improved, how the measurements changed and what still doesn't work well. We don't hide weaknesses.

Current state

98.2 %
Regular documents
95.6 %
Hard documents
98.2 %
Real public documents
92.9 %
Blind test

Rules + local model, 433 test documents in 17 languages, 0 false hits on texts without personal data. Measured 2026-10-03.

  1. v1.33 October 2026 · web tool + Chrome extension

    Our own local AI model, known names and a confidence light

    • Our own AI model (~53 MB, 17 languages) finds names, nicknames, addresses and companies the rules don't know. Bundled in the extension and works offline; in the web tool it downloads once when switched on and stays in your browser. Replaces the previous optional AI layer.
    • Known names: colleagues, clients and projects are always hidden — in any form, in capitals and without accents. CSV import; organisations can distribute the list through a Chrome policy.
    • Confidence light: doubtful words appear in a “Check” row — one click keeps them in the text.
    • Detection: chats (speakers, nicknames, Cyrillic), declined addresses, small businesses without a legal form, form field labels, gaps found by the blind test.

    Measurements

    Regular documents96.7 % → 98.2 %
    Hard documents89.8 % → 95.6 %
    Real public forms and contracts (new)98.2 %
    Blind test — texts the engine had never seen (new)92.9 %
    False hits on texts without personal data0 → 0

    Known limitations

    • On unknown texts (blind test) roughly one item in 14 stays unmarked. Review the highlighted parts before sending.
    • Informal chats are the weakest area: nicknames, lowercase names, typos.
    • Some languages (English, French, Slovak) have only a few hard test documents, so their hard-document figure is not yet reliable.
  2. v1.2.12 October 2026 · Chrome extension

    Faster pasting in Gemini

    • Gemini: long text (37,000 characters) pastes in 0.2 s instead of 6.2 s; the cursor stays after the text.
    • Companies written as “D.O.O” without the final dot and “D.0.0.” from scans.
    • Terms of use and a “Support the project” link in the panel.
  3. v1.230 September 2026 · web tool + Chrome extension

    17 languages and real documents

    • New languages: Polish, Dutch, Romanian, Czech, Slovak, Hungarian; English for the UK and US. Language is detected per paragraph, labels follow the language of the text.
    • Extension interface in 14 languages.
    • Real documents (public contracts and forms): domestic account numbers, IBANs across line breaks, health insurance numbers, phones, forms, layout-aware PDF reading; fewer false hits.
    • “Also hide amounts” switch; the side panel opens by itself when cleaning.

    Measurements

    Test documents136 → 263
    Regular documents96.5 % → 96.7 %
    Hard documents89.1 % → 89.8 %
    False hits on texts without personal data1 → 0
  4. v1.128 September 2026 · web tool + Chrome extension

    First public measurements

    • “How well it works” section: share of personal data found on test documents the engine had not seen, and false hits — re-checked after every change.
    • Spanish and Portuguese, passwords and keys, server names, references, tables.

    Measurements

    Regular documents (136)96.5 %
    Hard documents (28)89.1 %
  5. v1.027 September 2026 · web tool + Chrome extension

    First version

    • Local cleaning of personal data in the browser: declined names, consistent labels, real data restored in the AI's answer.
    • IDs with checksums (national IDs, tax numbers, IBAN …) for Slovenia, the Balkans, Germany, Austria, Italy and France.
    • PDF and image export with redacted areas.
    • Chrome extension: ChatGPT, Claude, Gemini, Copilot, Perplexity, DeepSeek, Mistral, Grok and more; real data only in the side panel.