Skip to main content
Favicon of LangSmith

LangSmith

What is LangSmith?

LangSmith is an AI observability and evaluation platform for teams that need to trace, debug, score, and deploy agents. It combines Observability, Evaluation, Deployment, Prompt Hub, Playground, and Canvas, plus annotation queues for human feedback. LangSmith works with Python, Typescript, Go, or Java SDKs and is used by Klarna, Vanta, monday.com, and Cloudflare. Plans run Developer $0 / seat per month, Plus $39 / seat per month, and Enterprise custom pricing.

Last verifiedHow we evaluate

Screenshot of LangSmith website

At a glance

Best for
LangSmith is best for AI teams who need to trace, evaluate, and ship agents with production visibility.
Pricing
Developer $0 / seat /mo; Plus $39 / seat /mo; Enterprise Custom

What the name covers now

If you last looked at LangSmith as "the LangChain tracing tool", your mental model is about a year out of date, and it will mislead you when you read pricing or comparisons.

In the changelog week of 13–17 October 2025, LangChain folded LangGraph Platform into LangSmith as LangSmith Deployment, and LangGraph Studio became LangSmith Studio. The changelog states that existing deployments, APIs, workflows, pricing and contracts were unchanged and no action was required. The Series B announcement five days later described LangSmith as "a comprehensive platform for agent engineering" spanning Observability, Evaluation, Deployment and a no-code Agent Builder.

As of today the pricing page bills under the LangSmith name for six distinct things: Observability & Evaluation, Deployment, Fleet, Engine, Sandboxes and an LLM Gateway. They have separate meters and separate free allowances.

The practical consequence: a blog post, price quote or comparison article about "LangSmith" written before late 2025 is about the tracing product only, and so is most of the community commentary quoted below. Check what is being priced before you compare anything.

You do not need LangChain to use it — and this is the most common mistake buyers make

The most persistent belief about LangSmith is that it drags LangChain in with it. It is worth addressing directly, because it is the reason a lot of teams never evaluate it, and the evidence says it is no longer true.

The belief is real and recent. A commenter on Hacker News in August 2025 put it as sharply as anyone: "LangSmith, the platform for prompt authoring/experimentation/observability isn't great but is useable… the fact that LangSmith more or less requires you to use Langchain to some degree is probably the biggest knock against it." The team at NonBios wrote that they "chose not to proceed with Langsmith as it seemed strongly coupled with Langchain."

The documentation and other practitioners contradict it. LangSmith exposes a standard OpenTelemetry OTLP endpoint, driven by two environment variables:

OTEL_EXPORTER_OTLP_ENDPOINT=https://api.smith.langchain.com/otel
OTEL_EXPORTER_OTLP_HEADERS="x-api-key=<your langsmith api key>"

Self-hosted instances use your own base URL with /api/v1/otel appended. Because it is ordinary OTLP, a collector can fan the same spans out to LangSmith and another backend at once — the docs describe this explicitly, which makes a trial genuinely reversible.

Independent practitioners confirm it in practice. Hamel Husain, an ML consultant with no disclosed vendor tie, wrote: "I like LangSmith — it doesn't require that you use LangChain and is intuitive and easy to use." Ian Bicking reported setting "up Langsmith with my own project quite quickly… very similar to setting up Langfuse, activated with a wrapper around the OpenAI library."

The integrations page lists 37 named integrations, and the striking thing is how many belong to competing stacks: AutoGen, CrewAI, Google ADK, Mastra, Microsoft Agent Framework, OpenAI Agents, PydanticAI, Semantic Kernel, Strands Agents and the Vercel AI SDK, plus coding agents including Claude Code, OpenAI Codex, Cursor, OpenCode and VS Code Copilot, and voice stacks including LiveKit and Pipecat.

If framework lock-in was your reason for ruling LangSmith out, that reason no longer holds. It is fair to say LangChain has not managed to shift the perception.

What tracing costs — and a ten-fold discrepancy to resolve before you sign

Both of the vendor's own current sources describe the same two-tier retention model, and they disagree about what the base tier costs.

What both agree on. Traces come in two tiers: base, retained 14 days, and extended, retained 400 days (customisable per workspace on Enterprise). Developer ($0/seat, one seat) includes up to 5,000 base traces a month, then pay-as-you-go. Plus ($39/seat/month, unlimited seats) includes 10,000. Enterprise is custom, invoiced annually.

Where they diverge. The pricing page prices traces in LangChain Storage Units at $1.00/LSU, and its rate tooltip reads 0.005 LSU/base trace and 0.0025 LSU/extended trace upgrade. That is $0.005 per base trace — $5.00 per thousand, with an extended trace totalling $0.0075.

The billing documentation instead states plainly: "The Base Charge for a trace is .05¢ per trace. We priced the upgrade such that an extended retention trace costs 10x the price of a base tier trace (.50¢ per trace) including both metrics. Thus, each upgrade costs .45¢." That is $0.0005 per base trace — $0.50 per thousand.

The gap is a factor of ten on the metric that dominates most bills. On a million traces a month beyond your allowance, the pricing page implies roughly $5,000 and the docs roughly $500.

This is not obviously a stale document. The docs repository shows src/langsmith/usage-and-billing.mdx was last edited on 31 July 2026, two weeks before this listing, and revised repeatedly through June and July 2026.

Two further data points, neither conclusive. The pricing page's base rate ($0.005) is exactly the docs' extended rate, which is what a straight repricing would look like. And a practitioner writing on Hacker News in June 2025 cited "$0.50 per 1k traces", matching the docs figure — evidence that the docs number was current at least through mid-2025, not evidence of what is charged today.

What to do: if traces are a material line item, ask for the current per-trace base and extended rates in writing before you commit. Any vendor will answer that. Do not build a forecast from either published number alone, and treat the per-1k prices quoted in third-party comparison articles as stale.

Retention auto-upgrades, and how to keep them from surprising you

Independent of the rate dispute above, trace billing has a mechanic that catches teams out.

An extended trace costs ten times a base trace, and certain actions promote a base trace to extended. Per the docs, these do:

  • Online evaluators scoring the trace, when the evaluator's retention setting is enabled
  • Automation rules with retention extension enabled matching any run in the trace
  • Feedback via API or SDK explicitly passing extend_trace_retention=true (extendTraceRetention: true in TypeScript)
  • Experiments, whose runs are created at extended retention by default

These do not: feedback and notes submitted through the LangSmith UI, and manually adding runs or threads to an annotation queue.

The part to note is the default: "Retention extension is enabled by default for new online evaluators and automation rules." You can opt out per evaluator or per rule, and for an evaluator the opt-out is only offered when the project's default retention is the base tier. So a team switching on online evaluation across a busy project without changing that default has moved those traces to the 10x rate by default rather than by choice.

To the vendor's credit, none of this is buried. The docs flag it with a warning box, set out the reasoning — the argument is that traces someone actually interacted with are worth more and should cost more — and document the controls: workspace spend limits that LangSmith converts into base and extended trace limits, a slider to distribute limits across tiers, per-project default retention, and automation rules that extend only a sampled percentage of traces. Datasets are retained indefinitely, so promoting traces is the wrong tool for long-term data collection.

The 14-day base window is itself a live complaint. One practitioner's summary: "They charge you even more if you want to keep a trace for more than 14 days, and even then you can only keep traces for a max of 400 days, even if you are hosting on your own servers." Both halves of that check out against the vendor's own documentation.

One side effect to plan for: once the extended-trace limit is reached, LangSmith blocks the actions that create more of them, so automation rules and evaluators that extend retention stop running. Tracing itself continues.

A dated repricing for Deployment: 1 October 2026

If you run agents on LangSmith Deployment, there is a migration with a hard date that appears in the docs and not on the pricing page.

Deployment is now billed on resources consumed rather than per run: compute in LangChain Compute Units (LCU, $1.50 each) and database storage in LangChain Storage Units (LSU, $1.00 each). Published rates are 0.045 LCU per vCPU-hour and 0.006 LCU per GiB-hour for runtime, and 0.177 LSU per vCPU-hour and 0.025 LSU per GiB-hour for the database. The $1.50 LCU price is corroborated in the changelog, which notes you can buy "as little as $4.50 in Gateway Credits, equivalent to 3 LCUs".

The docs then state: "This usage-based model replaces the previous per-run and uptime pricing. Existing customers remain on their current pricing until October 1, 2026, then move to the new model."

Two things follow. If you are an existing Deployment customer, your bill changes on that date and the direction depends on your workload — always-on Dedicated deployments consume compute continuously, whereas Serverless can scale to zero after inactivity. And scale-to-zero, the main lever for cutting idle cost, is available only on the new pricing, so staying on legacy pricing until October means going without it.

Plus includes one free small Serverless deployment. Scale to zero is marked beta and the inactivity window is described as subject to change while the feature rolls out.

Ingestion reliability, which is the part that matters here

For an observability tool, the question is not whether the dashboard loads — it is whether traces arrive. LangChain publishes enough to answer that, though not easily.

The status page is at global.status.smith.langchain.com (the older status.smith.langchain.com redirects there; status.langchain.com does not resolve). It is an incident.io page split into four regions — AWS US, GCP APAC, GCP EU, GCP US — tracking LangSmith Application, API, Run Ingestion, Deployments Control Plane, Deployments Data Plane, Billing, PromptHub, Fleet, Sandboxes and Bulk Exports separately. It is fully client-rendered, so incident history is not retrievable without a browser.

On the vendor's own published figures for May–August 2026, most components sit at 100% and Run Ingestion at 99.94% — the only component below 100%, and the one that matters most.

A third-party monitor, StatusGator, has tracked the page since November 2024 and lists dated incidents including: partial ingestion failures for traces with large payloads (6 August 2026), agent tool calls failing on Anthropic Claude models (5 August 2026), high latency and elevated 500s on analytical queries (3 August 2026), delayed ingestion (12 July 2026), increased API backend latency (13 May 2026), and login plus run ingestion impact (14 November 2025).

Read that proportionately. 99.94% on ingestion is a good number, and every one of these was resolved in hours. But three incidents fell in the first week of August 2026 and two touched the ingestion path, so if you intend to depend on LangSmith traces for production alerting rather than debugging, ask about ingestion SLAs specifically and design for delayed rather than guaranteed-instant traces. We are quoting StatusGator's dated incident list, not its aggregate outage count, which counts every degraded-status event and reads far worse than the uptime the vendor publishes for the same period.

Self-hosting: what it requires, and what still leaves your network

Self-hosting is real and well documented, with three topologies, but two constraints decide whether it is available to you.

It is an Enterprise add-on. The docs describe self-hosted LangSmith as "an add-on to the Enterprise plan designed for our largest, most security-conscious customers", and separately state that self-hosted deployments "require an Enterprise plan and the LangSmith license key delivered with that plan". The pricing page's hosting row confirms it: Developer and Plus are Cloud only; Enterprise is "Cloud, Hybrid, or Self-Hosted". RBAC is likewise Enterprise-only. There is no self-serve self-hosted tier.

The three topologies are the full self-hosted platform (control plane UI and APIs, observability, evaluation and deployment management inside your network), hybrid (LangChain-hosted control plane, your data plane), and standalone Agent Servers with your own PostgreSQL and Redis and no control plane — the lightest option, and the one intended for air-gapped use.

Egress is required unless you hold an offline licence. Unless you run air-gapped, a self-hosted instance must reach https://beacon.langchain.com. Three flows use it, and they are not equally optional:

  • Billing telemetry — licence verification and subscription/usage reporting. Required; the docs say it "cannot be disabled". Running with no egress means asking your account team for an offline (air-gapped) licence.
  • Operational telemetry — logs, metrics and traces for support diagnostics. On by default as of version 0.11, disabled in langsmith_config.yaml.
  • Usage telemetry — anonymised usage snapshots. On by default, disabled with PHONE_HOME_USAGE_REPORTING_ENABLED: false.

LangChain states it does not collect payload contents, database records, PII or anything identifying end users, and filters logs to error severity only.

Engine is the exception that matters for regulated buyers. Engine's orchestration runs in your VPC, but its model work does not: it sends what it needs to LangSmith Intelligence, a LangChain-managed zero-data-retention service, over a second egress path. The docs say plainly that each request carries "the trace content, code, and intermediate outputs Engine needs" and that traces "can include user messages, tool outputs, and PII". LSI does not persist prompt or response content and retains model and token metadata for billing. Air-gapped instances cannot run Engine at all; every other LangSmith feature works offline.

That candour in a self-hosting doc is unusual and genuinely useful — it answers the gating question directly instead of making you ask sales.

A migration already in flight: SmithDB, removal 31 January 2027

LangChain is moving LangSmith's trace query layer onto a new store called SmithDB. If you call the LangSmith API or SDK directly, this is work on your roadmap rather than a background detail.

The legacy endpoints — the v1 runs query and retrieve endpoints, v1 run sharing and public-run read endpoints, POST /api/v1/datasets/{dataset_id}/runs, and the annotation queue run endpoints — are deprecated. Published timeline:

DeploymentDeprecatedRemoved
All Cloud regionsEnd of July 202631 January 2027
Self-hostedv0.16v0.18

The SmithDB-backed methods need minimum SDK versions: Python langsmith >= 0.10.15, TypeScript langsmith >= 0.8.9, Java langsmith-java 0.1.0-beta.22, Go langsmith-go v0.25.4, CLI langsmith-cli v0.2.44. On self-hosted, the deprecated methods stop working once ClickHouse is disabled.

This is being handled well rather than badly, and that is worth saying: deprecated endpoints return Deprecation: true and a Sunset date header with a Link to the migration guide, the guide is written to be pasted into a coding agent to perform the migration, and it sits under a published policy giving Cloud six months from announcement to removal. Users of the UI and the standard tracing SDKs are unaffected — this concerns direct API and query-method consumers.

Operational ceilings worth checking against your throughput

The rate limits are published in full, which makes it possible to check fit before signing rather than after.

At the load balancer, on every plan: 5,000 requests per minute for POST/PATCH on /runs* and the same for /feedbacks*, 2,000 per minute across everything else, and 30 per minute for GET /runs/:id and for DELETE /sessions*. The SDK batches up to 100 runs per call, which keeps most applications clear of these.

Per-plan hourly ceilings, where an event is a run creation or update — a run created then updated in the same hour counts twice:

PlanEvents/hourData ingest/hour
Developer (no payment on file)50,000500 MB
Developer (payment on file)250,0002.5 GB
Startup/Plus500,0005.0 GB
EnterpriseCustomCustom

Developer accounts with no payment method are additionally capped at 5,000 unique traces a month. A single trace is limited to 25,000 runs, after which further runs on that trace are rejected — worth knowing if you run long-lived agents or deep recursive graphs.

Exceeding any of these returns HTTP 429. LangChain recommends retry with exponential backoff and jitter, and notes candidly that if you saturate endpoints for extended periods, retries will not save you.

Company, compliance and signs of life

Funding. LangChain announced a $125M Series B at a $1.25B valuation on 20 October 2025, led by IVP. This is one of the few things here corroborated well beyond the vendor: Fortune broke it as an exclusive on the day, with TechCrunch and SiliconANGLE reporting the same figures. Existing investors Sequoia, Benchmark and Amplify participated alongside new investors CapitalG and Sapphire Ventures, with strategic participation including ServiceNow Ventures, Workday Ventures, Cisco Investments, Datadog and Databricks. Earlier rounds were a $10M seed led by Benchmark in April 2023 and a $25M Series A led by Sequoia at a $200M valuation. We found no later round in reputable press, which is a negative search result rather than proof none exists.

On revenue, Fortune noted a July 2025 TechCrunch report putting ARR between $12M and $16M; a LangChain spokesperson called that "low for where we are today" and confirmed the company is not profitable. Treat it as a disputed estimate, not a figure. Separately, aggregator totals of "$260M raised" do not reconcile with the three announced rounds, which sum to $160M — prefer the announcements.

Compliance. The trust centre lists SOC 2 Type II, ISO 27001:2022, GDPR and HIPAA, with a 2025 SOC 2 Type 2 audit report, a 2026 ISO 27001 certificate and statement of applicability, a 2026 LangSmith penetration-test executive summary and a pre-filled CAIQ v4.0.3 questionnaire available on access request; the docs add that a BAA is available on request. That is a complete enterprise evidence pack, and the documents being gated behind an access request is standard practice. It is worth stating plainly that every one of these is LangChain asserting its own compliance posture — the underlying attestations are not public, so nothing here has been confirmed by a third party for this listing. The date LangSmith first achieved SOC 2 Type II is not established; the original changelog entry no longer resolves.

Maintenance. These are strong and trivially verifiable. The langsmith-sdk repository is MIT-licensed, was last pushed on 15 August 2026 — the day this listing was written — and shipped v0.11.0 on 14 August, v0.10.18 on 11 August and v0.10.17 on 7 August. The Python package on PyPI is at 0.11.0 under MIT. The changelog is published weekly and current through the week of 3–10 August 2026. There are SDKs for Python, TypeScript, Java and Go plus a CLI. Whatever else is uncertain here, nobody is asking whether this product is abandoned.

One clarification on "open source". The MIT licence covers the client SDK, not the product. LangSmith itself is closed-source hosted software, and running it on your own infrastructure requires an Enterprise contract and a licence key. LangChain and LangGraph — the frameworks — are separately open source, which is the usual source of the confusion.

Where it fits, and who should look elsewhere

It fits well if you want one system spanning tracing, evaluation and agent hosting rather than assembling three; if you need an enterprise evidence pack today; or if you want managed agent infrastructure with durable execution and human-in-the-loop, where Deployment has a real head start over general-purpose observability tools. The trace tree is the consistently praised part — one practitioner describes the value as "seeing full traces of moving through the graph and being able to inspect the inputs and outputs for each step."

Look harder if cost predictability at high trace volume is your main constraint. The metering is granular — LCUs, LSUs, seats, trace tiers, Gateway Credits, per-product allowances — and the discrepancy above means you should not commit to a volume forecast without written rates. If you are self-hosting to satisfy a no-egress requirement, confirm the offline licence path before designing around it. And if the people who need to read the output are not engineers, weigh that: the recurring criticism is that the strength is technical tracing rather than the visualisations or the packaged insights, and that evaluation features feel secondary to observability, with UI slowdowns on larger datasets.

A warning about the comparison articles you are about to read. We audited the search results for "LangSmith vs" and the great majority are published by direct competitors — Langfuse, Braintrust, Latitude, Confident AI, Laminar — or by SEO content mills with no original testing, one of which states LangSmith is "cloud-only", which the vendor's own pricing page contradicts. LangChain publishes its own comparison too. At least one widely-cited critique of LangSmith relayed as customer research turns out to be written by a competing vendor in its own launch post. None of these are worthless, but read the masthead first.

On independent evidence. There is no independent benchmark of LangSmith at scale that we could find, and we report that as absence rather than as a mark against it — it is a normal state for developer tooling. The most balanced independent read available comes from Hamel Husain, who works across the category and declines to pick a winner: the vendors he encounters most are "Langsmith, Arize and Braintrust", with features "very similar". One further observable signal: of the fifteen most-discussed LangSmith stories on Hacker News, three are launches of open-source alternatives to it, and LangChain's own LangSmith launches draw almost no discussion there. Read that two ways at once. Being the product people build alternatives to is a position of category leadership; it also reflects genuine demand for a free, self-hostable option — one such author's stated motivation was simply that "LangSmith needs a cloud account to see my own traces". That is precisely the gap Enterprise-only self-hosting leaves open, and if it describes you, look at the open-source observability tools before concluding LangSmith is the only option.

How much does LangSmith cost?

PlanPriceWhat's included
Developer$0 / seat per month
  • Up to 5k base traces / mo, then pay-as-you-go
  • Tracing to debug agent execution
  • Online and offline evals
  • Prompt Hub, Playground and Canvas for auto improving prompts
  • Annotation queues for human feedback
  • Monitoring and alerting
  • 1 Fleet agent
  • Community support
  • 1 seat
Plus$39 / seat per month
  • Everything in the Developer plan, and:
  • Up to 10k base traces / mo, then pay-as-you-go
  • 1 dev-sized agent deployment included
  • Email support
  • Unlimited Fleet agents
  • Add unlimited seats
  • Up to 3 workspaces
EnterpriseCustom pricing
  • Everything in the Plus plan, and:
  • Alternative hosting options, including hybrid and self-hosted so data doesn't leave your VPC
  • Custom SSO and RBAC
  • Acccess to deployed engineering team
  • Support SLA
  • Team trainings & architectural guidance
  • Custom Fleet packages

Frequently asked questions

What is LangSmith?

LangSmith is an AI observability and evaluation platform for teams that need to trace, debug, score, and deploy agents. It combines Observability, Evaluation, Deployment, Prompt Hub, Playground, and Canvas, plus annotation queues for human feedback. LangSmith works with Python, Typescript, Go, or Java SDKs and is used by Klarna, Vanta, monday.com, and Cloudflare. Plans run Developer $0 / seat per month, Plus $39 / seat per month, and Enterprise custom pricing.

How much does LangSmith cost? Is it free?

LangSmith has a free plan, with paid tiers including Plus at $39 / seat per month, Enterprise at Custom pricing.

What is LangSmith used for? Who is it for?

LangSmith is used for Observability, Evaluation, and Deployment. It's built for AI platform teams, ML engineers, and Product teams shipping agents.

Does LangSmith have an API and what does it integrate with?

LangSmith doesn't publish a public API. It integrates with OpenTelemetry, A2A, MCP, Agent Protocol, Pagerduty, and 6 more.

Editor's read

Developer includes up to 5k base traces per month, while Plus raises that to 10k and adds unlimited seats. If your trace volume is near either ceiling, verify how quickly pay-as-you-go charges would begin.

Share:

Sponsored
Favicon

 

  
 

Explore other Agent Tools & Integrations

Favicon

 

  
  
Favicon

 

  
  
Favicon