Configuration
Bukka supports YAML configuration when the command line starts to get unwieldy.
Configuration File Format
Bukka uses YAML configuration files for project setup. Generate a template with:
python -m bukka init-config
Default Configuration
The default template looks like this:
# Bukka Project Configuration Template
# Project settings
project:
name: my_ml_project
dataset: data/train.csv
target: target_column
skip_venv: false
enable_mlflow: false
mlflow_tracking_uri: null
# Data processing settings
data:
backend: polars
train_size: 0.8
stratify: true
strata: null
# Problem specification
problem:
type: auto # or: binary_classification, multiclass_classification, regression, clustering
Configuration Sections
Project section
The project section holds the basic project fields:
Parameter |
Type |
Description |
Default |
|---|---|---|---|
|
string |
Project name and directory. |
Required |
|
string |
Path to the dataset file. |
Optional |
|
string |
Target column name. |
Optional |
|
boolean |
Skip virtual environment creation. |
|
|
boolean |
Enable MLflow tracking. |
|
|
string |
Custom MLflow tracking URI. |
|
Example:
project:
name: fraud_detection
dataset: transactions.csv
target: is_fraud
skip_venv: false
enable_mlflow: true
Data section
The data section controls dataset handling:
Parameter |
Type |
Description |
Default |
|---|---|---|---|
|
string |
Dataframe backend. |
|
|
float |
Train/test split ratio. |
|
|
boolean |
Enable stratified sampling. |
|
|
list |
Columns for stratification. |
|
Example:
data:
backend: pandas
train_size: 0.75
stratify: true
strata: [gender, age_group]
Problem section
The problem section tells Bukka how to treat the project:
Parameter |
Type |
Description |
Default |
|---|---|---|---|
|
string |
ML problem type. |
|
Valid problem types:
auto- Automatic detectionbinary_classification- Binary classificationmulticlass_classification- Multi-class classificationregression- Regressionclustering- Clustering
Example:
problem:
type: binary_classification
Command-line arguments
Every configuration option can also be passed on the command line.
Priority order
When both a config file and CLI arguments are provided:
Command-line arguments (highest priority)
Configuration file
Default values (lowest priority)
Example:
# Config file sets backend to 'polars', CLI overrides to 'pandas'
python -m bukka run --config config.yaml --backend pandas
Global options
python -m bukka [OPTIONS] COMMAND [ARGS]
Options:
--help- Show help message
Init-config command
Generate a configuration template:
python -m bukka init-config [--output FILE]
Options:
-o, --output FILE- Output path (default:bukka_config.yaml)
Run command
Create a Bukka project:
python -m bukka run [OPTIONS]
Configuration:
-c, --config FILE- Path to YAML configuration file
Project settings:
-n, --name NAME- Project name (required unless using –config)-d, --dataset FILE- Dataset file path-t, --target COLUMN- Target column name-sv, --skip-venv- Skip virtual environment creation--mlflow- Enable MLflow tracking--mlflow-tracking-uri URI- MLflow tracking URI
Data processing:
-b, --backend NAME- Dataframe backend (default: polars)--train-size RATIO- Train/test split ratio (default: 0.8)--stratify- Enable stratified split (default: True)--no-stratify- Disable stratified split--strata COLUMNS- Stratification columns
Problem specification:
-p, --problem-type TYPE- ML problem type (default: auto)
Configuration examples
Example 1: Basic Classification
project:
name: iris_classifier
dataset: iris.csv
target: species
The rest of the old examples have been left out on purpose. The CLI pages are the primary reference for day-to-day use, and the API details are in API Reference.
- data:
backend: polars train_size: 0.8
- problem:
type: multiclass_classification
Usage:
python -m bukka run --config iris_config.yaml
Example 2: Regression with Custom Split
project:
name: house_prices
dataset: housing.csv
target: price
data:
backend: pandas
train_size: 0.7
stratify: false
problem:
type: regression
Usage:
python -m bukka run --config housing_config.yaml
Example 3: Clustering with No Virtual Environment
project:
name: customer_segments
dataset: customers.csv
skip_venv: true
data:
backend: polars
problem:
type: clustering
Usage:
python -m bukka run --config clustering_config.yaml
Example 4: Stratified Sampling
project:
name: medical_prediction
dataset: patient_data.csv
target: diagnosis
data:
backend: polars
train_size: 0.8
stratify: true
strata: [gender, age_group, severity]
problem:
type: binary_classification
Usage:
python -m bukka run --config medical_config.yaml
Example 5: MLflow Integration
project:
name: tracked_experiment
dataset: experiment_data.csv
target: outcome
enable_mlflow: true
mlflow_tracking_uri: "http://localhost:5000"
data:
backend: polars
train_size: 0.8
problem:
type: binary_classification
Usage:
python -m bukka run --config mlflow_config.yaml
Configuration Validation
Bukka validates all configuration values:
Project Name Validation
Cannot be empty
Cannot contain invalid characters:
< > : " | ? *
# ❌ Invalid
python -m bukka run --name "my:project"
# Error: Project name contains invalid characters
Dataset Path Validation
File must exist
Path must point to a file (not a directory)
# ❌ Invalid
python -m bukka run --name proj --dataset nonexistent.csv
# Error: Dataset file not found
Backend Validation
Must be a supported backend
Supported backends (via Narwhals):
polars(default, recommended)pandaspyarrowmodincudfdask
# ❌ Invalid
python -m bukka run --name proj --backend invalid
# Error: Backend 'invalid' not supported
Problem Type Validation
Must be a recognized ML problem type
Valid types:
autobinary_classificationmulticlass_classificationregressionclustering
Train Size Validation
Must be between 0.0 and 1.0
# ❌ Invalid
python -m bukka run --name proj --train-size 1.5
# Error: train_size must be between 0 and 1
Environment Variables
Bukka respects standard Python environment variables:
PYTHONPATH- Additional module search pathsVIRTUAL_ENV- Current virtual environment path
Best Practices
Version Control: Commit your
bukka_config.yamlfiles to version control.Template Reuse: Create template configs for common project types.
Documentation: Add comments to your YAML files to document choices.
Validation: Test your config with
init-configbefore running.Overrides: Use CLI arguments for temporary overrides, config files for permanent settings.
Modularity: Create separate configs for different experiments.
Troubleshooting
Config File Not Found
Error: Configuration file not found: config.yaml
Solution: Check the file path is correct and the file exists.
Invalid YAML Syntax
Error: Invalid YAML syntax in config file
Solution: Validate your YAML with a linter or:
python -c "import yaml; yaml.safe_load(open('config.yaml'))"
Missing Required Fields
Error: Project name is required
Solution: Either provide --name or set project.name in the config file.
Next Steps
See Usage Examples for practical examples
Check Getting Started for installation
Explore API Reference for programmatic usage