1. Technical Deep Dive
  2. Datasets & Ingestion

​
Datasets & Ingestion Workflow

Mako Code treats datasets as first-class citizens. Whether you drag-and-drop a Parquet file or generate a new dataset through code, the system enforces a uniform contract around storage, schema introspection, and versioning.


​
Data Directory Layout

At startup (main.py::init_data_directories) the backend creates:

data/
└── datasets/         # Your uploaded or saved .parquet files
    └── <dataset>.parquet

Each dataset may have a context file and one or more code notebooks referencing it.

└── context.json      # Human-editable metadata (description, owner, tags)

The schema is not stored separately—Polars infers it on demand.


​
Importing Data

​
1. Drag-and-Drop (Front-End)

Component: DataImportModal.svelte

  1. Triggers on ⌘⇧I or via sidebar Import button.
  2. Implements handleDrop() → uploadFile().
  3. Uses fetchApi("/dataset", { method: "POST", body: FormData }).

​
2. Programmatic Save (Back-End)

Call mako.save inside any executed Python snippet:

import polars as pl
from functions import mako

df = pl.DataFrame({"a": [1, 2, 3]})
mako.save(df, "tiny.parquet")

The helper writes to data/datasets/tiny.parquet and returns a Path so you can chain further logic.


​
Previewing Data

Selecting a dataset in the sidebar triggers:

GET /api/dataset/<path>

The back-end response includes:

  • schema – list of [name, dtype] pairs.
  • num_rows – extracted via pl.scan_parquet(...).collect().height.
  • preview – first n rows (default 100).

DataframeView.svelte paginates additional rows by requesting offset + limit query parameters.


​
Dataset Context

Context is optional JSON saved alongside the dataset. Fields include:

{
  "description": "UK sales in 2023 Q1",
  "owner": "sarah@acme.io",
  "tags": ["finance", "quarterly-report"],
  "examples": ["SELECT * FROM sales WHERE city = 'London'"]
}
  • Save via POST /api/dataset/context (front-end function saveDatasetContext).
  • Retrieve via GET /api/dataset/context/{name} when opening dataset details.

The front-end allows inline editing; changes are persisted on blur.


​
Deleting Data

RightSidebar.svelte dispatches DELETE /api/dataset.

Safety nets:

  1. Backend verifies that path is inside DATASETS_DIR (no directory traversal).
  2. If a code tab currently references the dataset, the front-end warns before proceeding.

​
Performance Tips

  1. Lazy Scans – Always prefer pl.scan_parquet over pl.read_parquet for large files. This allows projection push-down.
  2. Column Pruning – In analysis scripts, .select(["col1", "col2"]) early in the pipeline saves memory.
  3. Partitioned Datasets – Mako Code currently treats each dataset as a single Parquet file. For big data, generate partitioned directories (ds/yyyy/mm/*.parquet) and update mako.save accordingly.

​
Roadmap

  • Remote data sources (S3, GCS) via fs-spec.
  • Interactive schema editing in the UI.
  • Data quality checks using pandera or pydantic-df.

​
Summary

Datasets in Mako Code are intentionally simple: Arrow-native Parquet files on a local disk. This minimalism enables ultra-fast scans, easy backup strategies, and straightforward debugging. Coupled with Polars’ performance, the ingestion pathway can handle millions of rows without external infrastructure.