RAG Cleaner

Paste text or HTML — or drop a .txt, .md, or .html file — and this page converts it to clean Markdown, then splits it into deterministic chunks. Download the result as Markdown, JSONL, or a ZIP bundle with a chunk manifest. Identical input always produces identical chunks, ids, and files.

Private by default: your text is processed entirely in this tab. It is never sent to a server, never logged, and never stored. This page makes no network requests with your content.

Clean & chunk

Auto-detect treats input that starts with < as HTML. HTML keeps headings, lists, tables, links, and emphasis as Markdown; script and style blocks are dropped.

Paste text or HTML, or use the file picker below. Files are read locally and never leave your device.

Choose a .txt, .md, or .html file, or drop one here

No file selected yet.

Chunks never exceed this size (50–8000).

Repeats the tail of each chunk at the start of the next (0–500).

Keeps sections together and splits on # headings.

Ready. Paste text or HTML, then press Clean & chunk.

When do you need to clean text or HTML before retrieval?

Raw text and HTML are rarely ready for an LLM or retrieval pipeline as-is. Four situations come up most often:

A worked example, step by step

Suppose you exported a short status page from a CMS and want to feed it into a retrieval pipeline. The raw snippet is 283 characters of HTML — including a tracking script that should never reach your vector store:

<h2>Project status</h2>
<p>We ship the <strong>onboarding flow</strong> on Friday. See the <a href="https://example.com/status">full status</a>.</p>
<script>track('page');</script>
<h2>Open questions</h2>
<ul>
  <li>QA sign-off by Wednesday</li>
  <li>Update the help docs</li>
</ul>
A <script> block, a link, inline emphasis, and two headings — all in 283 characters.
  1. Paste the snippet and press Clean & chunk. Auto-detect recognizes the leading <h2> as HTML, so the lenient tokenizer renders it to Markdown instead of treating it as plain text.
  2. Check the Markdown. The 283-character input becomes 182 characters of clean Markdown: the headings become ## Project status and ## Open questions, the bold text becomes **onboarding flow**, the link becomes [full status](https://example.com/status) — and the <script> block is dropped entirely.
  3. Chunk with heading-aware mode. At 200 characters maximum with “Start a new chunk at each heading” checked, the document splits into exactly 2 chunks, each beginning at a heading: chunk-bf5086fdb5d252da (112 chars, “## Project status …”) and chunk-bf5087fdb5d2548d (68 chars, “## Open questions …”).
  4. Run it again. The same input and options produce the same ids, text, and files every time. Download the JSONL or the ZIP bundle, and the manifest proves which settings produced which chunks.

How it works

The pipeline has three deterministic stages, all running in this tab:

  1. Normalize. Plain text is normalized (Unicode NFC, CRLF/CR to LF, trailing-whitespace removal, blank-line collapse). HTML is parsed with a lenient tokenizer and rendered to Markdown — headings, lists, tables, links, and emphasis are preserved, while scripts, styles, and other non-content blocks are dropped.
  2. Chunk. The Markdown is split into blocks and combined up to your maximum size. With heading-aware mode, a new chunk starts at every # heading. Oversized single blocks are split at word boundaries when possible.
  3. Stabilize. Every chunk gets a content-derived id (FNV-1a over the document id plus the chunk index) and a sequential index. The same input and options always produce the same ids, texts, and files — nothing random, nothing timestamped.

Downloads are generated with the browser's Blob API and never touch the network. The ZIP bundle contains one Markdown file per chunk, a full output.md, output.jsonl, and a manifest.json listing each chunk id, index, and size.

Supported inputs and limits

Frequently asked questions

Does my content leave my computer?

No. The page is a static bundle of JavaScript; normalization, chunking, and file generation all run in your browser tab. There is no server call that receives your text.

Why are the chunk ids always the same?

Chunk ids are derived from the content itself (a stable hash of the document id plus the chunk index), not from the clock or random numbers. That makes re-running, diffing, and deduplicating chunks reliable.

What is the manifest for?

manifest.json records the options you used, the document id, and every chunk's id, index, and character count — no content. It lets you prove which chunks came from which settings without re-processing.

Will tables and links survive conversion?

Yes. Tables become GFM pipe tables, links become [text](url), and images become ![alt](src). Cell content is escaped so a pipe inside a cell cannot break the table.

What chunk size should I use?

There is no single right answer: useful chunk sizes depend on your embedding model and your content. The tool defaults to 800 characters (range 50–8000), a reasonable starting point for many retrieval setups. Heading-aware mode helps when your document has meaningful sections, and overlap repeats context across boundaries. The honest advice is to try a few settings and measure retrieval quality on your own data — re-running with different options is cheap here because the output is deterministic.

Is this an AI tool?

No. RAG Cleaner performs deterministic text processing only: Unicode normalization, a lenient HTML-to-Markdown conversion, block-based chunking, and FNV-1a content hashes. There is no language model, no embeddings, no API calls, and nothing is sent anywhere. "RAG" describes the use case — preparing text for retrieval pipelines — not a model running on this page.

Part of Local Toolworks. Last reviewed: 2026-08-16.