← Back to all spotlights

Fixing RAG Ingestion with Unstructured: Smarter Document Parsing

Tired of broken tables and shredded PDFs in your RAG pipeline? Unstructured extracts clean, layout-aware elements from messy enterprise files.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

Every developer building a Retrieval-Augmented Generation (RAG) system goes through the exact same five stages of grief. Stage one is unbridled optimism: you spin up an embedding model, connect a vector database, and marvel as a chatbot summarises your tidy Markdown notes. Stage two is reality: someone dumps a 400-page scanned quarterly report into the intake folder.

Suddenly, naive recursive text splitters shred multi-column pages into alphabet soup. Financial tables get sliced directly down the operating profit line, headers mingle with footers, and your multi-thousand-pound reasoning model hallucinates wildly because its context window received half a sentence from page 12 stitched directly into a footnote from page 13.

The PDF format was invented in 1993 to ensure printed ink landed exactly where a designer wanted it on physical paper. It was never intended to convey semantic meaning to a transformer model. This is where Unstructured enters the picture to do the dirty, unglamorous plumbing work of modern AI.


What is Unstructured-IO/unstructured?

Unstructured is an open-source Python library designed to ingest, clean, partition, and transform raw, unstructured documents—including PDFs, Word documents, PowerPoint decks, HTML, and images—into structured, semantically coherent JSON elements tailored for vector stores and LLM pipelines.

Rather than treating a document as a giant string of flat characters, Unstructured identifies visual layouts, categorising blocks into discrete document elements such as Title, NarrativeText, ListItem, and Table.


Raw File (PDF, DOCX, PPTX, HTML)
             │
             ▼
   [ Partitioning Strategy ]
 (fast / ocr_only / hi_res)
             │
             ▼
  Structured Element Stream
 (Title, NarrativeText, Table)
             │
             ▼
[ Chunking & Metadata Enrichment ]
 (chunk_by_title, coordinate bounds)
             │
             ▼
 Downstream Vector DB / RAG Ingestion

How It Works: Layout-Aware Partitioning

The core problem with traditional text splitters (like the default character or token chunkers found in early LangChain tutorials) is context blindness. If an earnings table sits inside a two-column layout, standard text extractors read across the horizontal scan line, weaving Column A and Column B together into gibberish.

Unstructured tackles this with modular partitioning pipelines. When you run partition_pdf, you choose a strategy based on your compute budget:

1. Fast: Extracts raw text streams directly from the document object model. It is lightning-fast, zero-GPU, and great for digital native PDFs without complex layouts.

2. OCR Only: Uses Tesseract to pull text out of image scans when no digital text layer exists.

3. Hi-Res: Employs computer vision models (such as layout detection neural networks) to segment bounding boxes on the page, separate two-column text, isolate images, and extract tables with their HTML cell relationships intact.

Naive Splitting vs Element Partitioning

CapabilityNaive Character ChunkingUnstructured Partitioning
Table PreservationSlices rows into arbitrary token blocksExtracts tables as structured text/HTML
Document StructureDrops headings or merges them with body textIdentifies Title elements to preserve hierarchy
Multi-Column PDFsReads horizontally across columnsSegments columns independently via layout models
Metadata TaggingMinimal (usually just file name and index)Page numbers, bounding boxes, file type, parent IDs

Hands-on: Ingesting and Chunking a Messy Document

To get started locally, install the core library along with the PDF extras. Depending on your needs, you may also require system dependencies like poppler-utils and tesseract-ocr.


pip install "unstructured[pdf]"

Here is a practical example showing how to parse a complex document, isolate tables, and chunk the narrative text intelligently using Unstructured's layout-aware engine:


from unstructured.partition.pdf import partition_pdf
from unstructured.chunking.title import chunk_by_title

# 1. Partition the PDF with layout detection
elements = partition_pdf(
    filename="quarterly_report.pdf",
    strategy="hi_res",
    infer_table_structure=True,
    languages=["eng"]
)

# 2. Inspect extracted elements
for el in elements[:5]:
    print(f"[{el.category.upper()}]: {el.text[:60]}...")
    if el.category == "Table":
        # Extract preserved HTML structure for the vector store
        print(f"Table HTML: {el.metadata.text_as_html[:100]}...\n")

# 3. Create context-aware chunks for embedding
# Instead of splitting at raw character counts, chunk_by_title groups
# text sections under their respective headings.
chunks = chunk_by_title(
    elements,
    max_characters=1500,
    combine_text_under_n_chars=200,
    new_after_n_chars=1000
)

print(f"\nGenerated {len(chunks)} semantically coherent chunks.")

Notice the chunk_by_title function. Instead of blindly cutting off mid-sentence at character 1,000, it respects document taxonomy. It keeps sub-points attached to their preceding heading and prevents bulleted lists from drifting into unrelated sections.


The Real-World Nuance: Managing the Overhead

While Unstructured solves the layout dilemma, developers in community forums and YouTube teardowns frequently point out the trade-offs:

  • Heavyweight Dependencies: Running strategy="hi_res" pulls in deep learning frameworks and layout detection models. If you deploy this on a minimalist serverless container (like AWS Lambda), you will quickly run into package size and memory ceilings.
  • Processing Latency: Running local OCR and vision-based document segmentation over hundreds of pages takes real compute time. For production-scale batch pipelines, you will want to cache partitioned outputs into intermediate JSON formats or run ingestion workers asynchronously.
  • The Sweet Spot: Use strategy="auto". It inspects the file, determines whether extractable text exists, and only falls back to heavy vision models when it detects complex layouts or scanned pages.

Why It Belongs in Your AI Stack

The current enterprise consensus is unambiguous: retrieval quality makes or breaks an agentic workflow. Better rerankers and smarter LLMs cannot fix garbage data injected during the ingestion phase.

By treating documents as structural ecosystems of titles, paragraphs, and tables rather than dumb strings, Unstructured bridges the gap between chaotic real-world paperwork and pristine LLM vector representations. If your RAG retrieval still trips over PDFs, swapping your document parser for Unstructured is the highest-leverage architectural fix you can make this afternoon.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.