Skip to main content

vaex

Processes and analyzes tabular datasets too large for RAM using Vaex, a Python library for lazy, out-of-core DataFrames over memory-mapped HDF5 and Arrow files, with CSV and Parquet import/export. Covers virtual columns, filtering, groupby aggregations, large-data heatmaps and histograms, and vaex-ml transformers, PCA, and K-means. Use when opening or converting multi-gigabyte CSV/HDF5/Arrow/Parquet files, computing fast statistics on billions of rows, visualizing massive datasets, building ML pipelines that do not fit in memory, or speeding up slow aggregations with lazy evaluation and delay=True. Prefer polars when data fits in RAM, or dask for cluster-distributed work.
Category: data-science-and-ml · License: MIT · Version: 1.1

Install

When to use it

Processes and analyzes tabular datasets too large for RAM using Vaex, a Python library for lazy, out-of-core DataFrames over memory-mapped HDF5 and Arrow files, with CSV and Parquet import/export. Covers virtual columns, filtering, groupby aggregations, large-data heatmaps and histograms, and vaex-ml transformers, PCA, and K-means. Use when opening or converting multi-gigabyte CSV/HDF5/Arrow/Parquet files, computing fast statistics on billions of rows, visualizing massive datasets, building ML pipelines that do not fit in memory, or speeding up slow aggregations with lazy evaluation and delay=True. Prefer polars when data fits in RAM, or dask for cluster-distributed work.

Full playbook

Read SKILL.md for the complete workflow, references and any scripts. The agent installer copies the full skill folder.