Modal logo

Modal

Serverless compute platform for AI inference, fine-tuning, and batch jobs with sub-second cold starts

OverviewArticles

From the Modal blog

Published by Modal, not by us. Every card opens the original post.

Modal Clusters are generally available

Multi-node GPU clusters with RDMA, gang scheduled from Modal's shared capacity pool and billed by the second, behind a single decorator.

modal.com

Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engine

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

modal.com

How to serve trillions of tokens for trillion-parameter coding agents

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

modal.com

Product updates: Sandbox Sidecars, new models, a refreshed dashboard, and more

Recent product updates from Modal and news from around the community.

modal.com

Modal is expanding in Europe with our new London office

Modal is expanding, and hiring on all fronts across Europe.

modal.com

How Botika runs full-stack generative AI on Modal

Botika is building the next generation of tooling for agentic e-commerce brands, backed by a fleet of custom models trained and served on Modal.

modal.com

Qwen3.8-2.4T-A95B now available on Modal

Qwen3.8-2.4T-A95B by Alibaba, with a 1M token context window, is now available via Modal Auto Endpoints.

modal.com

Bringing serverless functions closer to the speed of wire

Modal’s Function Call data path is now >50ms faster. Our new routing layer is geographically distributed, so you can further reduce your network overhead.

modal.com

A note on the Hugging Face agent incident

Hugging Face published a technical timeline of a recent agent intrusion. Modal's platform and isolation were not compromised in this incident.

modal.com

Kimi K3 by Moonshot now available on Modal

Kimi K3, a 2.8 trillion parameter multimodal model by Moonshot, along with a custom-trained DFlash speculator, is now available on Modal.

modal.com

Devin Outposts on Modal

Devin, built by Cognition, is an AI software engineer: it plans, writes, tests, and ships code semi-autonomously. With Outposts, Devin can now run its work in Modal sandboxes.

modal.com

Scaling to 1 million concurrent sandboxes in seconds

How (and why) we built a scheduling system that can scale to 1 million concurrent sandboxes (per workspace) in seconds.

modal.com