- Technical Deep Dive
- Backend Code Reference
Technical Deep Dive
Backend Code Reference
Line-by-line commentary on key backend files: ingestion.py, utils.py, sql_parser.py, and main.py.
Backend Code Reference
This document walks through the most critical Python modules in backend/, explaining implementation choices, error handling, and extension points. Reading source is encouraged—every snippet is replicated verbatim from the current main branch.
1. functions/ingestion.py
Purpose
- Guarantee dataset directories exist (
ensure_dataset_dir). - Provide helper wrappers around Polars I/O so other modules do not duplicate path logic.
Source Walk-through
import os
from pathlib import Path
import polars as pl
DATASETS_DIR_ENV = "DATASETS_DIR"
DEFAULT_DATA_DIR = Path("data/datasets")
def ensure_dataset_dir() -> Path:
"""Return a pathlib.Path pointing to DATASETS_DIR, creating it if absent."""
target_dir = Path(os.getenv(DATASETS_DIR_ENV, DEFAULT_DATA_DIR))
target_dir.mkdir(parents=True, exist_ok=True)
return target_dir
Uses parents=True to create nested directories in one call. The function’s return value is reused across modules, ensuring consistent path derivation even under custom environment overrides.
2. functions/mako.py
from pathlib import Path
from typing import Union
import polars as pl
from .ingestion import ensure_dataset_dir
def save(df: Union[pl.DataFrame, pl.LazyFrame], filename: str):
"""Persist a DataFrame to Parquet inside DATASETS_DIR."""
target = ensure_dataset_dir() / filename
if isinstance(df, pl.LazyFrame):
df.collect().write_parquet(target)
else:
df.write_parquet(target)
return target
save is automatically injected into the execution environment (exec_env in main.py). Users can call mako.save(df, "clean.parquet") from any code snippet to materialize intermediary results.
3. functions/sql_parser.py
While too simplistic for production-grade parsing, the regex approach keeps dependencies to zero and covers the 90th percentile of ad-hoc queries. Because the function returns (table, filter, columns), you can extend it to parse WHERE clauses and push them down into Polars scans.
SELECT_RE = re.compile(r"select\s+(?P<cols>.+?)\s+from\s+(?P<table>\S+)", re.I)
Edge-case: quoted identifiers are not yet supported.
4. functions/utils.py
This is the module where unsafe footprints may creep in—hence the careful selection of allowed imports.
4.1 execute_sql
import datetime as _dt
from .sql_parser import parse_sql_code
def execute_sql(code: str) -> pl.DataFrame:
table, _, cols = parse_sql_code(code)
dataset_path = ensure_dataset_dir() / f"{table}.parquet"
lf = pl.scan_parquet(dataset_path)
if cols != ["*"]:
lf = lf.select(cols)
result = lf.collect()
print(f"[{_dt.datetime.now().isoformat()}] SQL executed: {code[:60]}…")
return result
Notes
printstatements flow tostdoutcaptured by the API.- No
WHEREorJOINlogic—pull requests welcome.
4.2 delete_function
def delete_function(name: str) -> tuple[bool, str]:
path = Path("data/functions") / f"{name}.py"
if not path.exists():
return False, "Function not found"
path.unlink()
return True, ""
Front-end calls this via DELETE /api/functions/{name}. Because user helpers live outside site-packages, removal is trivial.
5. main.py
This is both the FastAPI app and execution sandbox.
5.1 Code Safety
validate_code_safety uses ast.walk to reject potentially dangerous nodes.
ILLEGAL_NODES = {ast.Import, ast.ImportFrom, ast.Call}
@classmethod
def validate_code_safety(cls, v: str) -> str:
tree = ast.parse(v)
for node in ast.walk(tree):
if isinstance(node, (ast.Import, ast.ImportFrom)):
raise ValueError("Import statements are disallowed")
return v
5.2 Initialization
init_data_directories is invoked at startup (@app.on_event("startup")). It guarantees all required directories exist—even when the container’s volume is empty.
for path in ["data", "data/datasets", "data/functions", "data/versions"]:
Path(path).mkdir(parents=True, exist_ok=True)
5.3 Streaming Logs (Planned)
A StreamingResponse variant of /api/execute is sketched in comments. Contributors familiar with FastAPI background tasks can pick this up.
6. Extending the Backend
- Add Arrow IPC: Replace
df.write_parquetwithpyarrow.ipc.RecordBatchFileWriterfor cross-language sharing. - SQL++: Leverage sqlglot for robust SQL → Polars translation.
- Notebook-Free DAG: Wrap code snippets in a directed acyclic graph (DAG) and cache node results.
Summary
The backend is intentionally minimal — fewer than 800 lines of Python across all modules — but powerful enough for day-to-day analytics. Its design ethos favors readability and hackability, making it an ideal starting point for advanced workflows.