Choose theme
AI Resources
liteparse
liteparse is a document parser from LlamaIndex that extracts text, Markdown and JSON, including PDF text positions for downstream retrieval and agent workflows.
Its native bindings handle PDFs, images and Office conversion. The browser WASM package has a narrower feature set, and choosing an external OCR engine changes where page images are processed. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Text with page positions
A parser whose PDF JSON output can connect extracted text to its location on the page. That supports source highlighting; it does not verify an answer's correctness.
Why it stands out
Separate OCR and layout signals
The complexity report distinguishes whether text needs OCR from whether the layout is complex. A digitally encoded page can still have difficult reading order.
Availability
Native bindings and browser WASM
Rust, Node and Python bindings accompany a browser package. Their supported formats, screenshot paths and OCR facilities differ.
Why it matters
What makes it useful
To highlight the source of a claim in a PDF, keep the extracted text coordinates. The visual-citations guide says they use PDF points with a top-left origin; a screenshot uses pixels. Its conversion scales positions by DPI divided by 72. Applying raw PDF coordinates to a 150-DPI screenshot would put the highlight in the wrong place.
What to know
Where it fits
Choose the runtime for the documents you have. Native Office parsing uses LibreOffice conversion. The browser guide instead describes PDF input as bytes and excludes Office conversion, page screenshots and built-in Tesseract OCR. A custom JavaScript OCR engine can be supplied there. Moving a native ingestion workflow into a browser therefore changes more than the package name. When exporting a report to Markdown, decide whether you also need its original embedded figures. The extraction and Markdown guides distinguish image references from decoded image files. Their explicit extraction route uses --extract-images with --image-output-dir to write those assets. Choose whether the report needs those figure files as well as Markdown references; a reference in the text is not itself the image bytes.
Notable points
What stands out
For routing pages, the complexity guide separates needs_ocr from layout.is_complex. A digital two-column page may need no OCR while still having a complex layout. The lit is-complex command's exit status signals OCR need, so a successful exit does not establish a simple reading layout. Use the layout field when that is the distinction your workflow needs.
Before using
What to review
For offline OCR, the OCR guide notes that Tesseract language files are fetched on first use. It describes preloading them through a supplied tessdata path or TESSDATA_PREFIX; installing the parser alone does not ensure those files are already cached.
A configured HTTP OCR server receives page images. The selected server determines where that processing happens, even when parsing starts on your own machine. Browser custom OCR can likewise call another service.
Reader fit
Who may find it relevant
If your service accepts PDFs from other people, the project's security guidance places responsibility for processing untrusted documents on the user. For that upload workflow, it recommends file validation, sandboxing, resource and time limits, and controlled concurrency. The extraction features alone do not supply those service controls.
Editorial note
Why LifeHubber lists it
We list liteparse for document builders who need to connect retrieved text to a specific region or table cell in the source. Its extraction guide describes extract_blocks output with bounding boxes for headings, paragraphs, tables and table cells. That supplies a way to inspect the surrounding region of a retrieved passage or the cell behind a table value, beyond converting text coordinates into screenshot pixels. Empty padding cells may have no bounding box, and a visible source region does not itself verify an answer.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Ecosystem
Keep browsing this category
Explore more AI ecosystem resources.
Laya
NandhaKishorM/laya
An early Apache-2.0 model family and Python runtime for bounded choice, score, and yes-or-no decisions, with English, multilingual, and task-specialized checkpoints plus a router that selects between them.
AnyJev
nokia-applied-research/AnyJev
An Apache-2.0 Python framework that turns existing Transformers or vLLM-served language models into bounded choice, yes-or-no, and score decisions, with levels for zero-label bias correction, calibration, and question-specific decision heads.
OpenAI Privacy Filter
openai/privacy-filter
A privacy-filtering model and local toolkit for detecting and masking personally identifiable information in text, positioned around high-throughput sanitization workflows.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.