Skip to content

Cluster simulator

kavier.sdk.cluster simulates a fixed-size GPU cluster that runs jobs of known duration under one of four scheduling policies: distributed-fcfs, distributed-backfill, consolidated-fcfs (default), or consolidated-backfill.

Per job, it reports queue wait, start and end time, runtime, energy, and goodput. Per cluster, it reports makespan, utilization, peak GPUs in use, and peak queue depth, plus a timeline of GPUs in use and queue depth.

There are two cluster goodput measures. goodput_jobs_per_s is the scheduling throughput in jobs/s. scheduling_goodput is the scheduling efficiency, sum(runtime_s) / sum(turnaround_s): training time over wall-clock time.

Python:

from kavier.sdk.cluster import schedule
r = schedule([{"submit_s": 0, "gpus": 8, "duration_s": 3600}],
             policy="consolidated-fcfs", num_nodes=2, node_gpus=8)
print(r.cluster.makespan_h, r.cluster.utilization)

CLI:

kavier cluster --jobs jobs.csv --policy consolidated-backfill --num-nodes 2 --node-gpus 8

Jobs CSV: submit_s,gpus,duration_s[,nodes,power_w_per_gpu].