This research is 65 days old. No newer filing has landed, but check the primary sources before acting on a number.
Founded in 2022 by Vipul Ved Prakash (previously founder of Topsy, sold to Apple), alongside Stanford's Percy Liang and Ce Zhang, Together AI set out to build the infrastructure layer open-source AI didn't have: a cloud purpose-built for training, fine-tuning, and serving open models at a fraction of closed-lab pricing. The thesis was contrarian when they started. It reads as consensus now that enterprises are voting with their token spend.
Research
The Together AI dossier
Researched July 7, 2026
The verdict
The best-run pure-play inference neocloud and the cleanest way to own the open-source-serving thesis — but it is a ~45%-gross-margin GPU reseller wearing a software cape, priced at 8x forward bookings into a market where the price of a token falls 90% by 2030; WATCHING, not a buy at $8.3B, until it proves the software layer (ATLAS/FlashAttention) actually widens the moat instead of just keeping pace.
Full research
Phase A — Understand the business
Company Overview
Together AI (legal entity Together Computer Inc.) is a pure-play AI cloud built around one bet: that the center of gravity in AI moves from a handful of closed frontier APIs (OpenAI, Anthropic) toward open-weight models (Llama, DeepSeek, Qwen/GLM, Kimi, MiniMax, Nemotron) — and that whoever serves those models fastest and cheapest owns a durable slice of the inference economy.
The business is two revenue lines stacked on the same GPUs:
Per-token API / serverless inference + fine-tuning — customers hit Together's endpoints for 200+ open models across chat, code, image, audio, vision, embeddings; billed per million tokens. Estimated ~30–40% of revenue.
GPU capacity rental (Together GPU Clusters / Instant Clusters) — renting Nvidia H100/H200/GB200 capacity by the GPU-hour for training, fine-tuning and dedicated serving, increasingly from Together's own data centers. Estimated ~60–70% of revenue.
Pricing, published: HGX H100 $1.76–$2.39/GPU-hr, H200 $3.15–$3.79, B200 $4.00–$5.50 (range = commitment vs on-demand); Batch API at a 50% discount to real-time.
Customers (named): Salesforce, Zoom, SK Telecom, The Washington Post, Zomato, Cursor, Cognition (Devin), Decagon, ElevenLabs, Suno, Pika, Krea, Hedra, Cartesia, Arc Institute; 450,000+ registered developers claimed. This is a healthy mix of AI-native scale-ups (the volume) and a handful of logo enterprises (the credibility).
Contract structure: mostly consumption / pay-as-you-go on the API side (spiky, low-commitment) and 1-month-minimum-to-multi-year reservations on clusters. This matters: the API book is not take-or-pay — it re-prices with every model and every competitor price cut. The cluster book carries some duration but far less than CoreWeave's multi-year hyperscaler contracts.
Founded June 2022. HQ San Francisco.
Supply Chain
Map upstream → Together → end customer, every named stakeholder:
Upstream (the chokepoints):
Nvidia — the master node. Supplies GPUs (H100/H200/GB200 NVL72), NVLink/NVSwitch, and the CUDA software moat. Nvidia is also an equity investor (Series A, B, and C) — a supplier that owns a piece of the customer. Single-source for the compute substrate.
HBM + advanced packaging — the true physical bottleneck. HBM3E supply (SK Hynix, Samsung, Micron) and TSMC CoWoS packaging capacity gate Blackwell availability industry-wide. Together does not control any of this.
Data-center / build partners — Hypertec Cloud (co-building a 36,000-GPU GB200 NVL72 cluster announced Nov 2024); and, per the Series C, 500+ MW of compute capacity to be "capitalized independently by investors" — i.e. Together is not putting the buildings on its own balance sheet. Historically also sourced from Crusoe Cloud, Vultr and 10+ GPU platforms.
Power — increasingly the binding constraint for the whole neocloud sector ("GPU race → power wars").
Together (the value-add layer): proprietary inference/training software — FlashAttention-4, Together Megakernel, together.compile, the ATLAS speculator, custom CUDA kernels, scheduling/virtualization. This is the only part of the chain Together actually owns end-to-end.
Downstream: AI-native app companies (Cursor, Cognition, Pika, Suno, ElevenLabs) and enterprises (Salesforce, Zoom, WaPo). These buyers are themselves highly price-sensitive and multi-homed — they will route to whoever is cheapest/fastest per token.
Chokepoint verdict: Together sits in the thinnest part of a chain it does not control at either end — Nvidia above (supply + a cut), commoditizing app buyers below. Its leverage is entirely in the software middle.
Competitive Advantages (moats)
What is genuinely defensible:
A research-grade software stack, publicly proven. Chief Scientist Tri Dao authored FlashAttention / FlashAttention-2/3/4 — used by OpenAI, Anthropic, Meta, Mistral. ATLAS adaptive speculative decoding reports ~500 tok/s on DeepSeek-V3.1 vs 105 on FP8 baseline (~4x) and up to 400% faster than vLLM; coding workloads 31% more TPS than the next-fastest OSS engine. This is a real performance edge — the question is whether it is durable (see Lens 13).
Full-stack integration — owning kernels → scheduling → serving lets Together extract more tokens per GPU-hour than a naive renter, which is how a ~45%-GM reseller stays alive against hyperscalers.
Open-source mindshare / distribution — RedPajama (30T-token dataset), the model zoo, LangChain/Vercel/MongoDB integrations, 450k developers = a genuine top-of-funnel.
Founder–researcher density — ~50% of staff are researchers; co-founders Percy Liang & Chris Ré are Stanford faculty.
What is NOT a moat:
The GPUs themselves (anyone can rent Nvidia).
Switching costs are low on the API side — it's an OpenAI-compatible endpoint; re-pointing a base URL to Fireworks/Baseten/DeepInfra is trivial.
Pricing power is capped: Together prices inference at roughly breakeven — below Fireworks, above loss-leaders.
Bargaining power:weak upstream (Nvidia dictates allocation and takes a revenue share on financed capacity), weak downstream (buyers are multi-homed and price-sensitive). The moat is real but narrow and software-shaped — it must be re-earned every model generation.
Segments
No audited segment disclosure exists (private, no Form 10-KA company’s audited annual report to the US regulator. The most complete thing it publishes.) — n/a — private, not disclosed for GAAP segment/geography splits.
Best available revenue-mix estimate:
Line
Est. share of revenue
Character
GPU rental / clusters
~60–70%
Capacity resale + software optimization; capital-intensive; some contract duration
API / serverless inference + fine-tuning
~30–40%
High-velocity, spiky, re-prices constantly; the "software company" part of the story
Geography: US-anchored (Maryland data center live Jul 2025) with international expansion (Sweden infrastructure live Sep 2025). No breakout.
Trend: the mix is the whole debate. Bulls want the API/software share to rise (higher-margin, stickier). Today's ~60/70% weighting toward raw rental is why the blended gross margin sits at ~45% — closer to a BMaaS reseller than a software company (see Lens 10/12).
Phase B — Measure performance
Funding & Valuation Trajectory (+private swap for "Earnings Result")
The private-company analogue of an earnings print is the round + the numbers disclosed around it.
Funding history:
Round
Date
Raised
Post-money valuation
Lead(s)
Source
Seed / early
2022–23
—
~$500M (implied at Series A)
Kleiner Perkins, Nvidia, Emergence
Series A
Nov 2023
$102.5M
~$500M
Kleiner Perkins
Series A ext.
Mar 2024
$106M
$1.25B
Salesforce Ventures, Coatue, Lux
Series B
Feb 2025
$305M
$3.3B
General Catalyst + Prosperity7
Series C
2026-07-01
$800M
$8.3B
Aramco Ventures
Series C participants: Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, March Capital, Pegatron, SE Ventures (SentinelOne/Schneider), Salesforce Ventures, DTCP Growth, Lux Capital, Geodesic, PSP Partners.
Total raised to date: ~$1.33B.
The revenue/traction numbers disclosed around the round (the "print"):
Annual bookings > $1.15B as of last quarter.
~$1.0B annualized revenue hit Feb 2026, up from ~$618M at end-2025 — i.e. ~62% growth in ~2 months of annualized run-rate, an extraordinary ramp.
~400% YoY revenue growth in 2024; $100M ARR reached in <10 months (by Sep 2024).
Valuation step-up: 2.5x in ~16 months ($3.3B → $8.3B).
Implied multiple: at $8.3B post on ~$1.15B bookings, ~7.2x bookings; on ~$1.0B annualized revenue, ~8.3x revenue. Not cheap, but below where CoreWeave and the frontier-model labs trade on forward revenue.
Balance-sheet flags (the private version): the buildout is being pushed off Together's balance sheet — the 500 MW is "capitalized independently by investors". Read charitably: capital-light, preserves equity. Read skeptically: it disperses the risk but also the control, and layers obligations (see Lens 13).
Founder & Ecosystem Signals (+private swap for "Earnings Calls")
No earnings calls (private). The analogue is founder interviews, product cadence, and ecosystem tone.
Product velocity is the loudest signal. In ~12 months Together shipped: Instant Clusters (self-serve GPU, GA Sep 2025), the 36k-GPU GB200 build (Hypertec), ATLAS, FlashAttention-4, Together Megakernel, together.compile, expanded post-training APIs. This is a team shipping at frontier-lab pace.
Consistent, falsifiable messaging: the thesis ("open-source inference is where the volume goes") is stated the same way across the Series B (Feb 2025) and Series C (Jul 2026) and is now backed by an external stat — open-model usage "tripled in twelve months". CEO Vipul Ved Prakash frames the raise entirely around scaling capacity 50x over five years.
What they've started saying: capacity/megawatts/power (the sector-wide pivot from "GPUs" to "power"). What's underplayed: unit economics and gross margin — never volunteered, always sourced by third parties. That silence is itself a tell.
Sentiment trend: unambiguously rising confidence, validated by a top-tier crossover/strategic syndicate. But confidence in a private raise is endogenous — Aramco marking you up is not the same as public-market price discovery.
Cap Table & Secondary Marks (+private swap for "Comps")
Syndicate quality — the IPO-proximity read:
Strategics stacked deep: Nvidia (3 rounds — supplier + owner), Salesforce Ventures (also a customer), Pegatron (ODM/hardware), SE Ventures (Schneider — power), and now Aramco Ventures leading (sovereign energy capital chasing compute).
Crossover / growth-equity present:Vista Equity Partners and General Catalyst and Coatue (prior) — the T. Rowe/Fidelity-type "IPO is in view" tell is partially there (Vista/Coatue) but not the classic mutual-fund crossover marks yet.
Read: this is a strategic-heavy, energy-tilted cap table — capital that wants compute exposure and offtake, not just financial return. Bullish for capacity access; it also means the company is being built as critical infrastructure for its investors, which can distort incentives.
Secondary marks / mutual-fund markups:n/a — not disclosed. No public secondary marks located.
Peer comps (the private/public inference-neocloud set — multiples `` or n/a):
CoreWeave is the scale reference (10x+ the revenue, multi-year hyperscaler backlog) but is a raw-capacity play; Together is the software-layer / open-model-serving play. They are adjacent, not identical — do not force a shared multiple. Fireworks/Baseten valuations could not be sourced → n/a rather than fabricate.
Traction & Funding Catalysts (+private — no public stock)
No listed stock, so no ±5% tape. The analogue: the events that re-rated the company in private markets.
Feb 2025 — $305M Series B @ $3.3B (2.6x step from $1.25B): validated the neocloud thesis pre-DeepSeek.
~Jan–Feb 2025 — DeepSeek moment: cheap, capable open models detonated demand for open-model serving — the single biggest tailwind to Together's core thesis.
2026 — ATLAS / FlashAttention-4 / Megakernel: the performance proof points that justify the "software, not reseller" narrative.
2026-07-01 — $800M Series C @ $8.3B: the re-rate, on >$1.15B bookings and open-model usage tripling.
Pattern: the market re-rates Together on (a) open-model secular adoption inflections (DeepSeek) and (b) capacity + software proof points — not on any single customer. Concentration risk on the catalyst side is low; the dependency is on the open-source thesis staying true.
Phase C — Judge people & books
Management
Vipul Ved Prakash (CEO, co-founder) — two prior exits: Cloudmark (email security → Proofpoint) and Topsy (social analytics, acquired by Apple ~$200M+); then senior AI/ML director at Apple 2013–2018. This is a proven repeat founder-operator, not a first-timer — the single strongest item in the file.
Tri Dao (Chief Scientist, co-founder) — FlashAttention author; arguably the most valuable individual technical asset in the inference-optimization world. His work is the credibility spine of the entire moat argument.
Ce Zhang (CTO, co-founder) — ex-professor (ETH Zürich / U. Chicago), scalable data systems.
Percy Liang & Chris Ré (co-founders) — Stanford faculty; Ré previously founded Lattice.io (→ Apple 2017). Elite research pedigree; ~50% of staff are researchers.
Capital-allocation history (as a private): the tell is that the buildout is kept off-balance-sheet (500 MW capitalized by investors) — a disciplined, equity-preserving choice vs. CoreWeave's debt-heavy model. Whether it's wise depends on the revenue-share terms (opaque).
Skin in the game: founders retain meaningful equity (early-stage; typical). Exact insider ownership n/a — not disclosed.
Red flags: the James v. Together Computer copyright suit ties directly to a founder-led research artifact (RedPajama/Books3) — see Lens 10. Otherwise no promotional behavior, no related-party pattern surfaced.
Archetype:founder-led, research-native, proven-exit CEO — close to the ideal profile for a technical infra company at this stage. This is the part of the thesis that is unambiguously strong.
Forensic Red Flags
No audited statements → forensic analysis is structurally limited; flags are model-level, not statement-level.
The gross-margin reality vs the software narrative. Blended GM ~45%. Sector reference: BMaaS gross margins run 55–65% before depreciation, and at those levels "the model has almost no margin of safety … if utilization slips below 80% returns flatline". Together at ~45% blended implies the reseller mix dominates the software mix. The core forensic question: is the ~45% pre- or post-depreciation, and how much of "revenue" is low-margin capacity pass-through? Not disclosed → the biggest unquantified risk in the file.
Bookings ≠ revenue. The headline is "$1.15B bookings", while recognized annualized revenue is ~$1.0B. Bookings can include multi-period cluster commitments; leaning on the larger number in the raise is a (mild, common) framing choice to watch.
Off-balance-sheet capacity = hidden leverage. 500 MW "capitalized independently by investors" moves Capital expenditureMoney spent on long-lived things — buildings, machines, servers — rather than on running costs. off Together's books but creates revenue-share / offtake obligations whose terms are undisclosed. This is the private analogue of operating-lease obfuscation.
Nvidia vendor-financing circularity. Nvidia now offers a "revenue-sharing and credit-support model" — GPUs without full capex, in exchange for "a recurring, usage-linked share of cloud revenue," split percentages undisclosed. If Together uses it, its true cost of capital/COGS is opaque and Nvidia is financing the demand for its own chips.
Regulatory findings (required sub-section):
SEC (EDGAR LR + AAER):0 findings — Together has no CIK (private, not an SEC filer); no EDGAR enforcement search is possible.
Litigation — MATERIAL:James v. Together Computer, filed 2025-11-06, N.D. Cal., a class action for direct and contributory copyright infringement over the RedPajama dataset's incorporation of Books3 (~196,640 books copied without authorization). A parallel suit targets Salesforce on the same RedPajama basis. This is a genuine, live legal overhang tied to Together's own flagship open-data artifact — the one enforcement item that matters.
Non-SEC agencies (FTC/DOJ/FDA/CFPB): web search surfaced no material agency enforcement actions against Together AI as of 2026-07-07.
Summary: No SEC/AAER findings (not a filer). One material civil copyright class action (James v. Together Computer, Nov 2025) is outstanding. All findings unaudited per public sources; verified via SEC EDGAR EFTS (no CIK), web search, and public litigation trackers as of 2026-07-07.
Phase D — Project & stress-test
IPO-Readiness & Path-to-Tradeable (+private swap for "Forward Projection")
No EPS model (private, no share economics disclosed) — the +private lens is when does this become tradeable, and what unlocks the S-1.
Where it sits on the private→public ladder:
Revenue scale: ✅ > $1B annualized / $1.15B bookings — comfortably past the ~$500M-ARR threshold IPO-quality infra companies clear.
Growth: ✅ triple-digit, decelerating from the 400% base but still hyper-growth.
Syndicate: ✅ growth-equity + strategics (Vista, General Catalyst, Aramco) — the kind of book that pre-stages an IPO.
Governance / audited financials: ❓ not evidenced publicly — the gating item.
Profitability / margin proof: ❌ ~45% GM and undisclosed net economics; public markets will demand the unit-economics story Together currently withholds.
CoreWeave precedent: a comparable neocloud is already public (2025 IPO), so the path is paved and the comp set exists.
Estimated IPO window:2027–2028. A Series C at this scale usually presages one more late/crossover round or a direct IPO within ~18–36 months; the binding constraints are margin disclosure and audited governance, not scale.
The milestone that unlocks the S-1: demonstrating that the software layer lifts blended gross margin toward 55–60%+ as the API/inference mix grows — i.e. proving it is a software-margin business, not a capacity reseller. Until that print exists, an IPO would price on the reseller comp, not the software comp.
Brier forecast: not logged (--watchlist unattended; skip per SKILL — only log on a genuinely committed base case). Candidate for a future pass: "Together AI files an S-1 / direct-lists by 2028-12-31, p≈0.55."
Write-back: no research/private-watch.json entry exists for together-ai; our model ledger not updated here (out of wave scope — no watchlist/index edits). Flag for a follow-up: add together-ai to private-watch.json with stage: series-c, ipo_readiness: high, catalyst: margin-disclosure/S-1, dossier: companies/together-ai/the previous dossier.
Bull vs Bear
Bull case. Together is the best-run pure-play on the single most durable trend in AI infra — the migration of inference volume to open-weight models (usage tripled in 12 months). It pairs a proven repeat-founder CEO with the person who wrote FlashAttention, and the software stack (ATLAS 4x, FA-4, Megakernel) is demonstrably faster than open-source baselines — which is exactly what lets a ~45%-GM operator survive against hyperscalers by squeezing more tokens per GPU-hour. Revenue went $618M → $1B annualized in ~2 months; the Series C syndicate (Aramco energy capital + Vista growth equity + Nvidia + Pegatron) hands it capacity, power, and hardware access most rivals can't get, off its own balance sheet. If the API/software mix rises and margins follow, the reseller comp (~7x bookings) re-rates to a software comp and the $8.3B mark looks early. Bull target: a $15–20B private mark / IPO within 24 months if margins inflect.
Bear case (2–3 permanent-impairment risks).
Commoditization eats the token. Inference prices fell >40x from 2023→2025 and Gartner projects >90% further cost decline by 2030. Together prices at breakeven already. In a race to the bottom on an undifferentiated OpenAI-compatible endpoint with near-zero switching costs, the software edge has to keep winning every generation just to hold ~45% — a permanent treadmill.
Squeezed between Nvidia and the hyperscalers. Nvidia takes a cut on financed capacity (circular financing) above; AWS Bedrock now hosts both Anthropic and OpenAI plus open models and can cross-subsidize below. Together owns neither the chip nor the customer relationship at the enterprise tier.
The margin might never inflect. If the mix stays ~60/70% raw rental, Together is a BMaaS reseller with almost no margin of safety (returns flatline below 80% utilization) — priced at 7x bookings.
Pre-mortem (18 months out, thesis broke): DeepSeek-class model releases slow OR a hyperscaler undercuts open-model serving to near-zero; utilization dips below 80% as the 500 MW comes online into softening demand; the vendor-financing/off-balance-sheet obligations turn a growth story into a fixed-cost trap; the copyright suit adds a headline tax. The $8.3B mark gets frozen or marked down at the next round.
Are the multiples too high? ~7–8x bookings is defensible for the growth but only if you believe the software-margin story. Priced as a reseller, it's rich. This is the crux.
Contrarian view (what the market refuses to see): the bull consensus treats Together as a software company that happens to own GPUs. The financials say it's a ~45%-gross-margin GPU reseller with an excellent software team — and the market is paying a software multiple for a reseller margin. The tell will be the first disclosed gross-margin trend: if it's not climbing toward 55%+, the re-rate reverses.
Devil's Advocate (short-seller)
I am dismantling the bull case.
What structurally breaks the money machine: the product Together sells (tokens off open models) is deflating 90% by 2030, sold through a fungible, OpenAI-compatible endpoint. There is no lock-in — a customer re-points a base URL in an afternoon. The "moat" is a performance lead measured in months, in a field (vLLM, SGLang, TensorRT-LLM, every rival's kernel team) that copies fast.
Concentration: less customer-concentrated than feared, but existentially Nvidia-concentrated — supply, software (CUDA), and a revenue-share creditor. If Nvidia tightens allocation or shifts terms, Together has no alternative substrate at scale.
The most dangerous competitor bulls underestimate: not CoreWeave — the hyperscalers. AWS/GCP/Azure can serve open models at a loss to keep the account and cross-sell, and Bedrock already carries Anthropic + OpenAI + open weights. A specialist with ~45% GM cannot win a subsidy war.
Worst capital/structure move:off-balance-sheet 500 MW + possible Nvidia vendor financing — stacking obligations "in multiple directions … narrows the margin for error". It flatters today's capital-lightness and hides tomorrow's fixed cost.
Accounting/disclosure red flag: leaning on $1.15B bookings while never disclosing gross-margin direction or net burn. Undisclosed metrics in a hype raise are undisclosed for a reason.
What must hold for the $8.3B price: (1) open-model usage keeps compounding; (2) Together keeps a perpetual software-perf lead; (3) blended GM rises despite token deflation; (4) utilization stays >80% as 500 MW lands. All four must hold. Break any one and the mark is stale.
If growth disappoints 20–30%: at a reseller's margin structure, a growth miss + a utilization dip is not a haircut — it's a down round, because the entire valuation rests on the slope, not the level.
Single permanent-impairment scenario, and plausibility: a hyperscaler-led price collapse in open-model serving (Bedrock/Vertex at cost) commoditizes the API layer while Together is committed to 500 MW of fixed capacity. Plausibility: medium — it's the base-rate outcome of every compute-commoditization cycle; the only defense is the software lead staying ahead, which is exactly the unprovable variable.
Management Questions (ordered by information value)
What is your blended gross margin today, and what has its trend been over the last four quarters — split API/inference vs GPU-rental? (the single number the whole valuation rests on)
Of "$1.15B annual bookings," how much is recognized recurring revenue vs multi-period cluster commitments, and what is your net revenue retention?
Are you using Nvidia's revenue-share / credit-support financing, and if so what share of cloud revenue does Nvidia take — how does it affect your effective COGS?
On the 500 MW "capitalized independently by investors," what are the offtake / revenue-share obligations, and what utilization do you need to break even on it?
Your API endpoint is OpenAI-compatible with near-zero switching cost — what, concretely, stops a customer leaving for a cheaper token tomorrow, and what is your churn?
Token prices are projected down 90% by 2030. How does your revenue-per-token and margin-per-token survive that, and what's the volume elasticity you're banking on?
How do you win against AWS Bedrock / Vertex when they can serve the same open models at a loss to hold the enterprise account?
How durable is the ATLAS/FlashAttention performance lead — how many months before vLLM/SGLang/TensorRT-LLM close it, and how do you re-open it each generation?
What is your current cash burn and runway, and does the Series C reach cash-flow breakeven or another raise?
What % of revenue is your top 10 customers (AI-native scale-ups like Cursor/Cognition/Suno), and how exposed are you if one insources or fails?
What is your realized GPU utilization, and what is your plan if it dips below the ~80% BMaaS break-even as new capacity lands?
What is the status and reserved exposure of James v. Together Computer (RedPajama/Books3), and does it change how you source training data?
What is your timeline to audited financials and an S-1 — is a 2027–2028 IPO the plan, or one more private round?
How dependent is your growth on the continued release cadence of frontier open models (DeepSeek/Qwen/Llama) that you don't control?
Nvidia is supplier, investor, and creditor — how do you manage that concentration, and what's your substrate plan (AMD, custom silicon) if allocation tightens?