Skip to main content
Favicon of Modal

Modal

What is Modal?

Modal is an AI-native container runtime for developers and ML teams that runs inference, training, and batch processing without managing infrastructure directly. It packages code in Python, adds programmable infra, elastic GPU scaling, and memory snapshotting, and pairs them with built-in storage and unified observability. Modal integrates with Slack, Weights and Biases, and TensorBoard, and is used by Substack, Lovable, and You.com. Plans run Starter $0, Team $250/month, and Enterprise custom.

Last verifiedHow we evaluate

Screenshot of Modal website

At a glance

Best for
Modal is best for developers who need to ship AI workloads without managing infrastructure.
Pricing
Starter $0; Team $250; Enterprise Custom

What it does well

Modal lets you decorate ordinary Python functions to run on remote CPUs or GPUs (@app.function(gpu="H100")) without managing servers, containers, or an orchestrator — Modal builds the image, provisions the machine, and autoscales to zero. It supports fractional, per-second billing on CPU/memory (documented at $0.00003942/core/sec and $0.00000667/GiB/sec) rather than the hourly increments common on traditional clouds, and Modal's own worked example claims this saved money on a bursty 24-hour workload ($4,740 vs. $5,400) (https://modal.com/pricing). The SDK ships fast: eleven Python client releases in 2026 alone, with recent additions including a Sandbox filesystem API, RBAC, and a new low-latency @app.server() primitive for HTTP apps (https://modal.com/docs/reference/changelog). It also lists a real, audited SOC 2 Type 2 report and a documented HIPAA BAA path on Enterprise (https://modal.com/docs/guide/security).

GPU model support and what pricing you can and can't see up front

Modal supports a wide range of current GPUs — T4, L4, A10, L40S, A100 (40GB/80GB), H100, H200, B200, and B300 — selectable by string in code, with automatic fallback chains and up to 8 GPUs per container for most types (https://modal.com/docs/guide/gpu). What Modal does not publish is a rate card mapping each GPU type to a dollar figure: the public pricing page shows only CPU ($0.00003942/core/sec) and memory ($0.00000667/GiB/sec) rates, links GPU pricing back to itself in a circular reference ("GPU: See standard pricing" pointing at the same /pricing page), and gives only one aggregate example ($3.95/GPU-hr average across a mixed fleet) rather than a per-model number (https://modal.com/pricing). A buyer who wants to compare, say, H100 cost against RunPod or Lambda has to sign up and check the dashboard or run modal billing rates from the CLI to get an actual figure.

Plans, seats, and where costs step up

Starter is free to start, with $30/month of compute included, up to 3 seats, 100 containers, 10 GPU-concurrency, 1-day log retention, and no deployment rollbacks, custom domains, static IP, or SOC 2/HIPAA/SSO. Team is $250+/month compute, uncapped seats, $100/month included compute, 5,000 containers, 50 GPU-concurrency, 30-day logs, custom domains, static IP proxy, and SOC 2 compliance — but still no audit logs, RBAC, or SSO, which remain Enterprise-only (verified against the live tier-comparison table, https://modal.com/pricing, 2026-09-14). Two multipliers apply platform-wide regardless of plan: non-preemptible execution costs 3x base CPU/memory price, and region selection outside the default costs 1.15–1.75x base price.

Company and funding

Modal Labs was founded in 2021 by CEO Erik Bernhardsson (former Spotify/Better.com CTO); the team is based in New York, Stockholm, San Francisco, and — as of a September 2026 blog post — a new London office (https://modal.com/company, https://modal.com/blog/modal-europe-expansion). Modal raised an $87M Series B in September 2025 at a $1.1B post-money valuation, led by Lux Capital, bringing total raised to $111M (https://modal.com/blog/announcing-our-series-b). In February 2026, TechCrunch reported Modal was in talks for a new round at roughly a $2.5B valuation with General Catalyst reportedly leading, and cited Modal's ARR at approximately $50M at that time — the CEO denied actively fundraising and called the conversations general — but no closed round at that valuation has been publicly confirmed since (https://techcrunch.com/2026/02/11/ai-inference-startup-modal-labs-in-talks-to-raise-at-2-5b-valuation-sources-say/).

Open-source surface and repository health

The product itself — scheduler, container runtime, storage layer — is closed-source and available only as Modal's hosted service; there is no self-hosted or on-prem option. What is open-source is the client: modal-client (the Python SDK/CLI, Apache-2.0, 514 stars, actively pushed to as of 11 Sept 2026) and modal-examples (MIT, 1,268 stars, pushed 9 Sept 2026), both under active development with no archival banners (https://github.com/modal-labs/modal-client, https://github.com/modal-labs/modal-examples, checked 2026-09-14). PyPI's modal package is current at 1.5.5, released 28 Aug 2026, and no vulnerabilities are recorded against it in OSV's database as of this check (https://pypi.org/project/modal/, https://api.osv.dev/v1/query).

Data handling and retention

Modal's security page publishes a specific retention table: function inputs/outputs are kept up to 7 days then deleted; Server/Endpoint request-response payloads are not stored at all (proxied directly); logs are retained 1 day on Starter, 30 days on Team, and configurable on Enterprise; Volumes and Images persist until manually deleted. Modal Inference endpoints are stated to be zero-data-retention, terminating TLS at the edge and forwarding to containers without writing payloads to disk (https://modal.com/docs/guide/security).

How much does Modal cost?

PlanPriceWhat's included
Starter$0
  • Built for small teams and independent developers looking to level up.
  • $30 / month free credits
  • 3 workspace seats included
  • 100 containers + 10 GPU concurrency
  • Scheduled and Web Functions (limited)
  • Real-time metrics and logs
  • Region selection
Team$250
  • $100 / month free credits
  • Unlimited seats
  • 1000 containers + 50 GPU concurrency
  • Unlimited Scheduled and Web Functions
  • Custom domains
  • Static IP proxy
  • Deployment rollbacks
  • Volume-based discounts
  • Embedded ML engineering services
  • Support via private Slack
EnterpriseCustom
  • For organizations prioritizing security, support and everlasting confidence.
  • Volume-based discounts
  • Unlimited seats
  • Higher GPU concurrency
  • Embedded ML engineering services
  • Support via private Slack
  • Audit logs, Okta SSO and HIPAA

Frequently asked questions

What is Modal?

Modal is an AI-native container runtime for developers and ML teams that runs inference, training, and batch processing without managing infrastructure directly. It packages code in Python, adds programmable infra, elastic GPU scaling, and memory snapshotting, and pairs them with built-in storage and unified observability. Modal integrates with Slack, Weights and Biases, and TensorBoard, and is used by Substack, Lovable, and You.com. Plans run Starter $0, Team $250/month, and Enterprise custom.

How much does Modal cost? Is it free?

Modal has a free plan, with paid tiers including Team at $250, Enterprise at Custom.

What is Modal used for? Who is it for?

Modal is used for Programmable infra, Elastic GPU scaling, and Unified observability. It's built for ML engineers, Platform teams, and Developers building AI apps.

Does Modal have an API and what does it integrate with?

Modal doesn't publish a public API. It integrates with telemetry providers, Python, WebRTC, WebSocket, Weights and Biases, and 2 more.

Editor's read

Check the GPU concurrency ceiling on Starter and Team before committing. Starter includes 10 GPU concurrency, while Team raises that to 50; workloads that burst beyond those limits will need Enterprise for higher GPU concurrency.

Share:

Sponsored
Favicon

 

  
 

Explore other Agent Tools & Integrations

Favicon

 

  
  
Favicon

 

  
  
Favicon