LIFEHUBBER
Choose theme

AI Resources

liteparse

GitHub stars: 12.8K GitHub forks: 876 Declared license: Apache-2.0: Apache-2.0 Last pushed October 5, 2026: Pushed 3d ago
Stats from GitHub

liteparse is a document parser from LlamaIndex that extracts text, Markdown and JSON, including PDF text positions for downstream retrieval and agent workflows.

Its native bindings handle PDFs, images and Office conversion. The browser WASM package has a narrower feature set, and choosing an external OCR engine changes where page images are processed. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Text with page positions

A parser whose PDF JSON output can connect extracted text to its location on the page. That supports source highlighting; it does not verify an answer's correctness.

Why it stands out

Separate OCR and layout signals

The complexity report distinguishes whether text needs OCR from whether the layout is complex. A digitally encoded page can still have difficult reading order.

Availability

Native bindings and browser WASM

Rust, Node and Python bindings accompany a browser package. Their supported formats, screenshot paths and OCR facilities differ.

Why it matters

What makes it useful

To highlight the source of a claim in a PDF, keep the extracted text coordinates. The visual-citations guide says they use PDF points with a top-left origin; a screenshot uses pixels. Its conversion scales positions by DPI divided by 72. Applying raw PDF coordinates to a 150-DPI screenshot would put the highlight in the wrong place.

Notable points

What stands out

For routing pages, the complexity guide separates needs_ocr from layout.is_complex. A digital two-column page may need no OCR while still having a complex layout. The lit is-complex command's exit status signals OCR need, so a successful exit does not establish a simple reading layout. Use the layout field when that is the distinction your workflow needs.

Before using

What to review

For offline OCR, the OCR guide notes that Tesseract language files are fetched on first use. It describes preloading them through a supplied tessdata path or TESSDATA_PREFIX; installing the parser alone does not ensure those files are already cached.

A configured HTTP OCR server receives page images. The selected server determines where that processing happens, even when parsing starts on your own machine. Browser custom OCR can likewise call another service.

Reader fit

Who may find it relevant

If your service accepts PDFs from other people, the project's security guidance places responsibility for processing untrusted documents on the user. For that upload workflow, it recommends file validation, sandboxing, resource and time limits, and controlled concurrency. The extraction features alone do not supply those service controls.

Editorial note

Why LifeHubber lists it

We list liteparse for document builders who need to connect retrieved text to a specific region or table cell in the source. Its extraction guide describes extract_blocks output with bounding boxes for headings, paragraphs, tables and table cells. That supplies a way to inspect the surrounding region of a retrieved passage or the cell behind a table value, beyond converting text coordinates into screenshot pixels. Empty padding cells may have no bounding box, and a visible source region does not itself verify an answer.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving