
Managing Multiple Slurm Clusters with SkyPilot
If you have access to multiple Slurm clusters, managing them manually is painful. SkyPilot lets you use them as a single unified resource pool.
Guides, posts, and resources from SkyPilot.

If you have access to multiple Slurm clusters, managing them manually is painful. SkyPilot lets you use them as a single unified resource pool.

Agentic RL spends most of every step on inference, not training. Scale inference in RL with SkyPilot Job Groups and slime.

Kimi K3's weights are open. Here's how to serve the 2.8T-parameter model across your own GPU clusters with a single YAML.


One YAML, one endpoint to run and manage inference on your own GPU fleet.

SkyPilot runs untrusted, LLM-generated code in sandboxes on the Kubernetes clusters you already own. A single cluster sustains 50,000+ sandboxes, with multi-cluster support to go further. Individual launches in less than a second. All at up to one-tenth the cost of hosted alternatives, with your code and data never leaving your cloud.

Online reinforcement learning for LLMs breaks Slurm's batch scheduling model. We'll discuss why, and what can be done about it.

We ran hundreds of benchmarks to tune storage systems for distributed training so you don’t have to.

Introducing GPU Compass: One dashboard to browse, compare pricing, and launch across every GPU cloud.

With the SkyPilot Agent Skill, your AI coding agent can launch clusters, run training jobs and manage cloud resources across any infrastructure using natural language.

Coding agents working from code alone generate shallow hypotheses. Adding a research phase — arxiv papers, competing forks, other backends — produced 5 kernel fusions that made llama.cpp CPU inference 15% faster.

Karpathy's autoresearch runs one experiment at a time. We gave it access to our GPU infra and let it run experiments in parallel.

Get the latest stories, tips, and insights delivered to your inbox.