1. Technical Deep Dive
  2. Backend Code Reference

​
Backend Code Reference

This document walks through the most critical Python modules in backend/, explaining implementation choices, error handling, and extension points. Reading source is encouraged—every snippet is replicated verbatim from the current main branch.


​
1. functions/ingestion.py

​
Purpose

  1. Guarantee dataset directories exist (ensure_dataset_dir).
  2. Provide helper wrappers around Polars I/O so other modules do not duplicate path logic.

​
Source Walk-through

import os
from pathlib import Path
import polars as pl

DATASETS_DIR_ENV = "DATASETS_DIR"
DEFAULT_DATA_DIR = Path("data/datasets")

def ensure_dataset_dir() -> Path:
    """Return a pathlib.Path pointing to DATASETS_DIR, creating it if absent."""
    target_dir = Path(os.getenv(DATASETS_DIR_ENV, DEFAULT_DATA_DIR))
    target_dir.mkdir(parents=True, exist_ok=True)
    return target_dir

Uses parents=True to create nested directories in one call. The function’s return value is reused across modules, ensuring consistent path derivation even under custom environment overrides.


​
2. functions/mako.py

from pathlib import Path
from typing import Union
import polars as pl
from .ingestion import ensure_dataset_dir

def save(df: Union[pl.DataFrame, pl.LazyFrame], filename: str):
    """Persist a DataFrame to Parquet inside DATASETS_DIR."""
    target = ensure_dataset_dir() / filename
    if isinstance(df, pl.LazyFrame):
        df.collect().write_parquet(target)
    else:
        df.write_parquet(target)
    return target

save is automatically injected into the execution environment (exec_env in main.py). Users can call mako.save(df, "clean.parquet") from any code snippet to materialize intermediary results.


​
3. functions/sql_parser.py

While too simplistic for production-grade parsing, the regex approach keeps dependencies to zero and covers the 90th percentile of ad-hoc queries. Because the function returns (table, filter, columns), you can extend it to parse WHERE clauses and push them down into Polars scans.

SELECT_RE = re.compile(r"select\s+(?P<cols>.+?)\s+from\s+(?P<table>\S+)", re.I)

Edge-case: quoted identifiers are not yet supported.


​
4. functions/utils.py

This is the module where unsafe footprints may creep in—hence the careful selection of allowed imports.

​
4.1 execute_sql

import datetime as _dt
from .sql_parser import parse_sql_code

def execute_sql(code: str) -> pl.DataFrame:
    table, _, cols = parse_sql_code(code)
    dataset_path = ensure_dataset_dir() / f"{table}.parquet"
    lf = pl.scan_parquet(dataset_path)
    if cols != ["*"]:
        lf = lf.select(cols)
    result = lf.collect()
    print(f"[{_dt.datetime.now().isoformat()}] SQL executed: {code[:60]}…")
    return result

Notes

  • print statements flow to stdout captured by the API.
  • No WHERE or JOIN logic—pull requests welcome.

​
4.2 delete_function

def delete_function(name: str) -> tuple[bool, str]:
    path = Path("data/functions") / f"{name}.py"
    if not path.exists():
        return False, "Function not found"
    path.unlink()
    return True, ""

Front-end calls this via DELETE /api/functions/{name}. Because user helpers live outside site-packages, removal is trivial.


​
5. main.py

This is both the FastAPI app and execution sandbox.

​
5.1 Code Safety

validate_code_safety uses ast.walk to reject potentially dangerous nodes.

ILLEGAL_NODES = {ast.Import, ast.ImportFrom, ast.Call}

@classmethod
def validate_code_safety(cls, v: str) -> str:
    tree = ast.parse(v)
    for node in ast.walk(tree):
        if isinstance(node, (ast.Import, ast.ImportFrom)):
            raise ValueError("Import statements are disallowed")
    return v

​
5.2 Initialization

init_data_directories is invoked at startup (@app.on_event("startup")). It guarantees all required directories exist—even when the container’s volume is empty.

for path in ["data", "data/datasets", "data/functions", "data/versions"]:
    Path(path).mkdir(parents=True, exist_ok=True)

​
5.3 Streaming Logs (Planned)

A StreamingResponse variant of /api/execute is sketched in comments. Contributors familiar with FastAPI background tasks can pick this up.


​
6. Extending the Backend

  • Add Arrow IPC: Replace df.write_parquet with pyarrow.ipc.RecordBatchFileWriter for cross-language sharing.
  • SQL++: Leverage sqlglot for robust SQL → Polars translation.
  • Notebook-Free DAG: Wrap code snippets in a directed acyclic graph (DAG) and cache node results.

​
Summary

The backend is intentionally minimal — fewer than 800 lines of Python across all modules — but powerful enough for day-to-day analytics. Its design ethos favors readability and hackability, making it an ideal starting point for advanced workflows.