vaex
Processes and analyzes tabular datasets too large for RAM using Vaex, a Python library for lazy, out-of-core DataFrames over memory-mapped HDF5 and Arrow files, with CSV and Parquet import/export. Covers virtual columns, filtering, groupby aggregations, large-data heatmaps and histograms, and vaex-ml transformers, PCA, and K-means. Use when opening or converting multi-gigabyte CSV/HDF5/Arrow/Parquet files, computing fast statistics on billions of rows, visualizing massive datasets, building ML pipelines that do not fit in memory, or speeding up slow aggregations with lazy evaluation and delay=True. Prefer polars when data fits in RAM, or dask for cluster-distributed work.
- Version
- 1.1
- License
- MIT
- Compatibility
- Requires Python 3.10+ (3.12+ recommended with vaex 4.19.0). Install with uv pip install vaex. Optional s3fs/gcsfs/adlfs for cloud I/O.
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Write
- Edit
- Bash
- Grep
- Glob
Files
- skills/vaex/SKILL.md
- skills/vaex/references/core_dataframes.md
- skills/vaex/references/data_processing.md
- skills/vaex/references/io_operations.md
- skills/vaex/references/machine_learning.md
- skills/vaex/references/performance.md
- skills/vaex/references/visualization.md
Every link opens the file at its source, pinned to the revision this page describes.