liteparse
Local document and PDF parsing that returns spatial text with bounding boxes. Use for extracting text from PDFs, DOCX, Office files, and images; running OCR on scans; producing layout-preserved JSON for RAG; batch-ingesting folders of papers; or rendering pages to PNG for multimodal agents. Distinguishing capabilities are spatial text boxes, Markdown, page raster output, and local parsing with optional custom HTTP OCR.
- Version
- 1.6
- License
- Apache-2.0
- Compatibility
- Python 3.10+ with liteparse 2.15.0. LibreOffice required for Office formats; images convert natively. Tesseract is bundled but missing language data downloads on first use. Optional HTTP OCR requires network access and server-specific authentication.
Pinned to revision 68105dd992f1, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Write
- Edit
- Bash
Files
- skills/liteparse/SKILL.md
- skills/liteparse/SKILL_CN.md
- skills/liteparse/references/api_reference.md
- skills/liteparse/references/choosing_a_parser.md
- skills/liteparse/references/cli_reference.md
- skills/liteparse/references/ocr_and_formats.md
- skills/liteparse/references/output_formats.md
- skills/liteparse/scripts/batch_parse_dir.py
Every link opens the file at its source, pinned to the revision this page describes.