Usage¶
The CLI, the dashboard, and the Python API connect to one core engine.
CLI¶
The coastline command has five subcommands. Each documents its flags with --help.
| Subcommand | Use |
|---|---|
recommend-job |
Search and rank the best configurations for one job, or for a CSV of jobs. |
recommend-trace |
Recommend a configuration for every job in a fine-tuning trace. |
simulate |
Predict the throughput and power of one configuration. |
explain |
Explain a recommendation: the score breakdown and the feasibility verdict. |
utils |
tune a data-driven predictor on your own runs, trace-to-runs, and plot-trace. |
Simulate one configuration, then explain the recommendation for the same job:
coastline simulate --model mistral-7b-v0.1 --method lora --gpu-model NVIDIA-A100-SXM4-80GB \
--tokens 2048 --batch-size 8 --gpus-per-node 4 --nodes 1
coastline explain --model mistral-7b-v0.1 --method lora --gpu-model NVIDIA-A100-SXM4-80GB \
--tokens 2048 --batch-size 8
--predictor selects the throughput model and --feasibility the feasibility check; see
How a recommendation is made.
One job¶
recommend-job --config reads the workload, the policy, the predictors, and the search grid from a configuration
file, and returns the best configuration as JSON:
workload:
llm_model: mistral-7b-v0.1
fine_tuning_method: lora # full | lora | qlora | gptq-lora
tokens_per_sample: 2048 # sequence length
batch_size: 8 # per device
strategy:
name: multi_objective # or min_gpu, which uses no grid
preset: performance # performance (default) | balanced | energy (or energy-saver)
predictors:
performance: kavier # kavier | cache | intelligent | a data-driven model
energy: kavier_power
feasibility: autoconf # autoconf | rules | none
grid:
gpu_models: [NVIDIA-A100-SXM4-80GB]
batch_sizes: [4, 8, 16, 32, 64] # batch sizes to search
total_gpus: [1, 2, 4, 8, 16, 32] # GPU counts to search
coastline recommend-job --config experiment.yaml --output-dir runs/exp-01
The recommendation goes to runs/exp-01/recommendation.json; without --output-dir it is printed. In a clone of
the repository, coastline recommend-job with no flags runs the job declared in
config/coastline_functionality/experiment.yaml. coastline recommend-job --interactive starts a guided REPL.
Many jobs¶
From a clone, a CSV of workloads goes in and a CSV with one recommendation per row comes out:
coastline recommend-job --config config/batch_config.yaml \
--input config/coastline_functionality/sample_workloads.csv --output recommendations.csv
coastline recommend-trace --input config/coastline_functionality/sample_trace.csv --output recommended_trace.csv
The workloads CSV has the columns llm_model, fine_tuning_method, gpu_model, tokens_per_sample, and
batch_size, and optionally the job's layout, gpus_per_node and number_of_nodes. recommend-trace reads a
fine-tuning trace with the columns of sample_trace.csv, where metadata.batch_size is each job's total batch; a
row whose total does not split evenly over its GPUs is kept unchanged under the weighted goals. --goal sets the
goal (default performance; min_gpu keeps each job's total batch), and --visual also draws the cluster
timeline (needs the [plot] extra).
Infrastructure¶
In a clone, config/coastline_functionality/infrastructure.yaml declares the available infrastructure: 32
NVIDIA-A100-SXM4-80GB GPUs, distributed equally across 4 compute nodes. Without that file, Coastline assumes 64
GPUs, 8 per node. --cluster-gpus overrides the total, and the INFRASTRUCTURE_CONFIG and EXPERIMENT_CONFIG
environment variables point to other files.
Dashboard¶
coastline-ui
The dashboard runs at http://127.0.0.1:8000; COASTLINE_UI_HOST and COASTLINE_UI_PORT change the address. A
user defines a fine-tuning workload, requests a recommendation, and appends the job to the workload queue. In admin
mode, the sysadmin imports a CSV of jobs and runs the queue on the cluster. The playground compares the predictors
on one configuration, through simulation-based or cache-retrieval what-if analysis.
Python API¶
A runnable script: python docs/usage.py.
"""Show the Coastline Python API on one fine-tuning workload and on a batch.
`import coastline` has three entry points over one engine:
coastline(predictor=...) a configured recommender; returns Recommendation objects, best first
coastline.recommend(batch) a DataFrame, a list of dicts, or a dict in; a DataFrame out
coastline.recommend_csv(...) a CSV of workloads and a config file in; a CSV of recommendations out
A workload sets llm_model, fine_tuning_method, gpu_model, tokens_per_sample, and batch_size
(per device). The goal is performance (the default), balanced, energy, or min_gpu. Here, the three
entry points search the same grid and pick the same configuration. min_gpu searches no grid: it
keeps the job's total batch (batch_size on 1 GPU here, since the jobs set no gpus_per_node or
number_of_nodes) and returns the first feasible GPU count in 1, 2, 4, ... with that batch split over
the GPUs. The `coastline` CLI and the `coastline-ui` dashboard call the same engine.
Run with: python docs/usage.py
"""
import tempfile
from pathlib import Path
import pandas as pd
import yaml
import coastline
job = {
"llm_model": "mistral-7b-v0.1",
"fine_tuning_method": "lora",
"gpu_model": "NVIDIA-A100-SXM4-80GB",
"tokens_per_sample": 1024,
"batch_size": 16,
}
batch_sizes = [4, 8, 16, 32] # per-device batch sizes to search
gpu_counts = [1, 2, 4, 8] # GPU counts to search
# 1) One workload. Kavier predicts throughput and power; AutoConf checks feasibility (the default).
# Without a goal, the recommender ranks for performance.
recommender = coastline(predictor="kavier")
ranked = recommender.recommend(job, batch_sizes=batch_sizes, total_gpus=gpu_counts)
for rank, rec in enumerate(ranked[:3], start=1):
print(f"{rank}. {rec.total_gpus} GPUs, batch {rec.metadata['batch_size']}, {rec.predicted_throughput:.0f} tokens/s")
# 2) A batch of workloads: one row per workload, with the chosen configuration and its predictions.
# runtime_s and energy_wh cover dataset_size samples for the given number of epochs.
other = {**job, "llm_model": "granite-3.3-8b", "fine_tuning_method": "full", "tokens_per_sample": 4096, "batch_size": 4}
jobs = pd.DataFrame([job, other])
for goal in ("performance", "balanced", "energy", "min_gpu"):
picks = coastline.recommend(
jobs, goal=goal, predictor="kavier", batch_sizes=batch_sizes, max_gpus=8, dataset_size=50_000, epochs=1
)
print(f"\n{goal}:")
print(picks[["llm_model", "total_gpus", "batch_size", "throughput_tok_s", "energy_wh"]].to_string(index=False))
# 3) CSV in, CSV out. The config file sets the policy, the predictors, and the search grid.
tmp = Path(tempfile.mkdtemp())
config = {
"strategy": {"name": "multi_objective", "preset": "performance"},
"predictors": {"performance": "kavier", "energy": "kavier_power", "feasibility": "autoconf"},
"grid": {"batch_sizes": batch_sizes, "total_gpus": gpu_counts},
}
(tmp / "config.yaml").write_text(yaml.safe_dump(config))
jobs.to_csv(tmp / "jobs.csv", index=False)
coastline.recommend_csv(tmp / "config.yaml", tmp / "jobs.csv", tmp / "recommendations.csv")
print(f"\nrecommend_csv wrote {tmp / 'recommendations.csv'}:")
print(pd.read_csv(tmp / "recommendations.csv")[["llm_model", "recommended_total_gpus", "predicted_throughput"]])