- Technical Deep Dive
- Datasets & Ingestion
Technical Deep Dive
Datasets & Ingestion
Everything you need to know about importing, previewing, and persisting datasets in Mako Code.
Datasets & Ingestion Workflow
Mako Code treats datasets as first-class citizens. Whether you drag-and-drop a Parquet file or generate a new dataset through code, the system enforces a uniform contract around storage, schema introspection, and versioning.
Data Directory Layout
At startup (main.py::init_data_directories) the backend creates:
data/
└── datasets/ # Your uploaded or saved .parquet files
└── <dataset>.parquet
Each dataset may have a context file and one or more code notebooks referencing it.
└── context.json # Human-editable metadata (description, owner, tags)
The schema is not stored separately—Polars infers it on demand.
Importing Data
1. Drag-and-Drop (Front-End)
Component: DataImportModal.svelte
- Triggers on
⌘⇧Ior via sidebar Import button. - Implements
handleDrop()→uploadFile(). - Uses
fetchApi("/dataset", { method: "POST", body: FormData }).
2. Programmatic Save (Back-End)
Call mako.save inside any executed Python snippet:
import polars as pl
from functions import mako
df = pl.DataFrame({"a": [1, 2, 3]})
mako.save(df, "tiny.parquet")
The helper writes to data/datasets/tiny.parquet and returns a Path so you can chain further logic.
Previewing Data
Selecting a dataset in the sidebar triggers:
GET /api/dataset/<path>
The back-end response includes:
schema– list of[name, dtype]pairs.num_rows– extracted viapl.scan_parquet(...).collect().height.preview– first n rows (default 100).
DataframeView.svelte paginates additional rows by requesting offset + limit query parameters.
Dataset Context
Context is optional JSON saved alongside the dataset. Fields include:
{
"description": "UK sales in 2023 Q1",
"owner": "sarah@acme.io",
"tags": ["finance", "quarterly-report"],
"examples": ["SELECT * FROM sales WHERE city = 'London'"]
}
- Save via
POST /api/dataset/context(front-end functionsaveDatasetContext). - Retrieve via
GET /api/dataset/context/{name}when opening dataset details.
The front-end allows inline editing; changes are persisted on blur.
Deleting Data
RightSidebar.svelte dispatches DELETE /api/dataset.
Safety nets:
- Backend verifies that
pathis insideDATASETS_DIR(no directory traversal). - If a code tab currently references the dataset, the front-end warns before proceeding.
Performance Tips
- Lazy Scans – Always prefer
pl.scan_parquetoverpl.read_parquetfor large files. This allows projection push-down. - Column Pruning – In analysis scripts,
.select(["col1", "col2"])early in the pipeline saves memory. - Partitioned Datasets – Mako Code currently treats each dataset as a single Parquet file. For big data, generate partitioned directories (
ds/yyyy/mm/*.parquet) and updatemako.saveaccordingly.
Roadmap
- Remote data sources (S3, GCS) via
fs-spec. - Interactive schema editing in the UI.
- Data quality checks using
panderaorpydantic-df.
Summary
Datasets in Mako Code are intentionally simple: Arrow-native Parquet files on a local disk. This minimalism enables ultra-fast scans, easy backup strategies, and straightforward debugging. Coupled with Polars’ performance, the ingestion pathway can handle millions of rows without external infrastructure.