
Modal
Serverless compute platform for AI inference, fine-tuning, and batch jobs with sub-second cold starts
From the Modal blog
Published by Modal, not by us. Every card opens the original post.
Modal Clusters are generally available
Multi-node GPU clusters with RDMA, gang scheduled from Modal's shared capacity pool and billed by the second, behind a single decorator.
Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engine
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
How to serve trillions of tokens for trillion-parameter coding agents
Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.
Product updates: Sandbox Sidecars, new models, a refreshed dashboard, and more
Recent product updates from Modal and news from around the community.
Modal is expanding in Europe with our new London office
Modal is expanding, and hiring on all fronts across Europe.
How Botika runs full-stack generative AI on Modal
Botika is building the next generation of tooling for agentic e-commerce brands, backed by a fleet of custom models trained and served on Modal.
Qwen3.8-2.4T-A95B now available on Modal
Qwen3.8-2.4T-A95B by Alibaba, with a 1M token context window, is now available via Modal Auto Endpoints.
Bringing serverless functions closer to the speed of wire
Modal’s Function Call data path is now >50ms faster. Our new routing layer is geographically distributed, so you can further reduce your network overhead.
A note on the Hugging Face agent incident
Hugging Face published a technical timeline of a recent agent intrusion. Modal's platform and isolation were not compromised in this incident.
Kimi K3 by Moonshot now available on Modal
Kimi K3, a 2.8 trillion parameter multimodal model by Moonshot, along with a custom-trained DFlash speculator, is now available on Modal.
Devin Outposts on Modal
Devin, built by Cognition, is an AI software engineer: it plans, writes, tests, and ships code semi-autonomously. With Outposts, Devin can now run its work in Modal sandboxes.
Scaling to 1 million concurrent sandboxes in seconds
How (and why) we built a scheduling system that can scale to 1 million concurrent sandboxes (per workspace) in seconds.