RAG Cleaner
Paste text or HTML — or drop a .txt, .md, or
.html file — and this page converts it to clean Markdown,
then splits it into deterministic chunks. Download the result as
Markdown, JSONL, or a
ZIP bundle with a chunk manifest. Identical input always
produces identical chunks, ids, and files.
Clean & chunk
Result
When do you need to clean text or HTML before retrieval?
Raw text and HTML are rarely ready for an LLM or retrieval pipeline as-is. Four situations come up most often:
- Scraped or exported web pages for LLM and RAG pipelines. Crawls, CMS exports, and “save as HTML” copies carry navigation, scripts, styles, and tracking markup that add tokens, slow embedding, and dilute retrieval quality. Converting to clean Markdown keeps headings, lists, tables, and links while dropping the noise — the same preprocessing step behind the popular “html to markdown” converters, done locally.
- Chunking documents for embeddings and vector search. Retrieval quality depends on consistent, meaningful chunks. This page splits Markdown at headings when you ask it to, respects a maximum size, and can repeat overlap across boundaries — so chunks stay coherent instead of cutting sentences in half.
- Reproducible preprocessing and evaluation. Because chunk ids are derived from the content and options, the same input always produces the same ids, order, and files. That makes pipeline diffs, deduplication, and before/after evaluations trustworthy: run it twice and compare, or keep the manifest as proof of which settings produced which chunks.
- Privacy-sensitive documents. Pasting proprietary or personal text into a hosted converter means sending it somewhere. Here normalization, conversion, chunking, and ZIP generation all happen in this tab — the page makes no network request with your content.
A worked example, step by step
Suppose you exported a short status page from a CMS and want to feed it into a retrieval pipeline. The raw snippet is 283 characters of HTML — including a tracking script that should never reach your vector store:
<h2>Project status</h2>
<p>We ship the <strong>onboarding flow</strong> on Friday. See the <a href="https://example.com/status">full status</a>.</p>
<script>track('page');</script>
<h2>Open questions</h2>
<ul>
<li>QA sign-off by Wednesday</li>
<li>Update the help docs</li>
</ul> - Paste the snippet and press Clean & chunk. Auto-detect
recognizes the leading
<h2>as HTML, so the lenient tokenizer renders it to Markdown instead of treating it as plain text. - Check the Markdown. The 283-character input becomes
182 characters of clean Markdown: the headings become
## Project statusand## Open questions, the bold text becomes**onboarding flow**, the link becomes[full status](https://example.com/status)— and the<script>block is dropped entirely. - Chunk with heading-aware mode. At 200 characters maximum with
“Start a new chunk at each heading” checked, the document splits into exactly
2 chunks, each beginning at a heading:
chunk-bf5086fdb5d252da(112 chars, “## Project status …”) andchunk-bf5087fdb5d2548d(68 chars, “## Open questions …”). - Run it again. The same input and options produce the same ids, text, and files every time. Download the JSONL or the ZIP bundle, and the manifest proves which settings produced which chunks.
How it works
The pipeline has three deterministic stages, all running in this tab:
- Normalize. Plain text is normalized (Unicode NFC, CRLF/CR to LF, trailing-whitespace removal, blank-line collapse). HTML is parsed with a lenient tokenizer and rendered to Markdown — headings, lists, tables, links, and emphasis are preserved, while scripts, styles, and other non-content blocks are dropped.
- Chunk. The Markdown is split into blocks and
combined up to your maximum size. With heading-aware mode, a new
chunk starts at every
#heading. Oversized single blocks are split at word boundaries when possible. - Stabilize. Every chunk gets a content-derived id (FNV-1a over the document id plus the chunk index) and a sequential index. The same input and options always produce the same ids, texts, and files — nothing random, nothing timestamped.
Downloads are generated with the browser's Blob API and never touch
the network. The ZIP bundle contains one Markdown file per chunk, a
full output.md, output.jsonl, and a
manifest.json listing each chunk id, index, and size.
Supported inputs and limits
- Formats: plain text, Markdown (treated as text),
and HTML (including fragments).
.txt,.md, and.htmlfiles can be dropped or picked. - HTML coverage: headings, paragraphs, lists (including nested), tables, blockquotes, code blocks, links, images, and inline emphasis. Scripts, styles, SVG, and iframes are removed.
- Size: bounded by this browser tab's memory. Very large documents are chunked in memory and may take a moment.
- Limits: PDFs, DOCX, EPUB, and images are not parsed in this prototype. Markdown tables use GFM syntax; pipes in cell text are escaped. HTML entity references are decoded.
Frequently asked questions
Does my content leave my computer?
No. The page is a static bundle of JavaScript; normalization, chunking, and file generation all run in your browser tab. There is no server call that receives your text.
Why are the chunk ids always the same?
Chunk ids are derived from the content itself (a stable hash of the document id plus the chunk index), not from the clock or random numbers. That makes re-running, diffing, and deduplicating chunks reliable.
What is the manifest for?
manifest.json records the options you used, the document id, and every chunk's id, index, and character count — no content. It lets you prove which chunks came from which settings without re-processing.
Will tables and links survive conversion?
Yes. Tables become GFM pipe tables, links become [text](url), and images become . Cell content is escaped so a pipe inside a cell cannot break the table.
What chunk size should I use?
There is no single right answer: useful chunk sizes depend on your embedding model and your content. The tool defaults to 800 characters (range 50–8000), a reasonable starting point for many retrieval setups. Heading-aware mode helps when your document has meaningful sections, and overlap repeats context across boundaries. The honest advice is to try a few settings and measure retrieval quality on your own data — re-running with different options is cheap here because the output is deterministic.
Is this an AI tool?
No. RAG Cleaner performs deterministic text processing only: Unicode normalization, a lenient HTML-to-Markdown conversion, block-based chunking, and FNV-1a content hashes. There is no language model, no embeddings, no API calls, and nothing is sent anywhere. "RAG" describes the use case — preparing text for retrieval pipelines — not a model running on this page.
Related tools
- Date Calculator — count days and business days between dates, or add and subtract days.
- Text Hygiene — inspect invisible characters and improve writing.
- Payload Workbench — inspect and transform JSON and Base64 payloads.
- Document Hygiene — remove metadata from source documents before sharing.
- Passphrase Generator — create strong random passphrases locally from the EFF word list.
- Bingo Card Generator — make free printable bingo cards for 75-ball, 90-ball, or word games.
- Media Optimizer — resize and compress images in your browser.
- Time Coordinator — plan schedules across time zones with DST-aware local times.
- Subtitle Converter — convert SRT and VTT subtitle files locally.
- Volumetric Weight Calculator — calculate chargeable weight for courier and freight shipments.
- CBM Calculator — estimate cubic meters (CBM) and container fit for sea-freight shipments.
- Color Contrast Checker — check the WCAG contrast ratio between two colors with AA/AAA verdicts.
- Freight Class Calculator — estimate the NMFC freight class from pallet dimensions and weight (standard LTL density scale, classes 50–500).
Part of Local Toolworks. Last reviewed: 2026-08-16.