Reno, Nevada-based startup Positron AI, which makes chips for inference, i.e., for running already trained AI models, raised $875M at a $5B valuation. The round was raised in two tranches, with participation from NEA, Andra Capital, Atreides Management, Valor Equity Partners, SemiAnalysis Capital, and Jim Clark, founder of Silicon Graphics and Netscape. Just seven months ago, the company was valued at just over $1B, so the valuation has grown more than fourfold. Funds will go to the next generation of hardware, the Asimov chip and Titan system, an engineering data center, and scaling the existing Atlas system. The architectural essence of the bet is large volumes of mass memory LPDDR5X on the chip instead of scarce HBM.

image

What happened

The round is divided into two parts: $375M Series C at a pre-money valuation of $3.5B, where co-leads were NEA, Andra Capital, Atreides Management, Valor Equity Partners, and SemiAnalysis Capital (Dylan Patel's analytics), and a Series C-1 of up to $500M from NEA and Jim Clark. The baseline for the current growth was February 2026, when a $230M Series B valued the company at just over $1B. The majority of funds are directed to the silicon chip Asimov: its tape-out on the TSMC N3P process is planned for the end of 2026, mass production for the second half of 2027, and the memory volume on a single chip will be from 288 GB to 2304 GB. Alongside it, the Titan system of four to eight Asimov chips is being designed, intended for models of over 16 trillion parameters and a context of over 10 million tokens per node. Additionally, the company will build an engineering data center with a capacity of over 2 MW and scale up production of the current Atlas system, which is already deployed in more than 50 racks in Oracle Cloud Infrastructure; among the named customers are Parasail, Jump Trading, and i3d.net. In the press release, the company also claims to have achieved over 90% of available memory bandwidth in its systems.

Context

The shift that this round records is happening at the level of the entire AI infrastructure: the center of gravity of the market is moving from model training to inference, where the bottleneck becomes memory, not computational power in FLOPS. The traditional path through expensive HBM memory and its CoWoS packaging runs into scarce supply chains, and it is from this dependency that Positron is consciously moving away, building an architecture around mass memory LPDDR5X. For the same reason, according to the round's context, Nvidia reduced plans for Rubin Ultra from about 1 TB of HBM4E memory to 192 GB. Positron formulates its bet through the economics of "tokens per dollar" and "tokens per watt," and its funds are supported, in addition to the already named funds, by Qatar Investment Authority, Cisco Investments, and Hudson River Trading. For the capital market, this is a sign that inference has ceased to be a side application to training GPUs and has become an independent segment of silicon for which investors are ready to pay at an early stage.

Why this matters for the industry

For the industry, the capital market has confirmed with money that inference silicon has become an independent asset: raising financing for such hardware has become easier, and GPU buyers have received an additional argument in negotiations with suppliers. The narrative that "memory is more important than FLOPS" is already influencing infrastructure planning, and the emergence of a significant non-GPU inference path diversifies the equipment market. If the bet on LPDDR5X is justified, by 2027 the commoditization of long context is possible: the cost of token generation may drop significantly, and memory will cease to be a scarce resource. The nearest checkpoint is the end of 2026: whether the Asimov tape-out on TSMC N3P will be confirmed on time or delayed. An expansion of Atlas deployments in Oracle Cloud Infrastructure and with third-party providers is expected, and the market will compare the company's promises with actual deliveries and revenue.

Why this matters for users

There is no direct effect on current models and services: Atlas serves specific customers through Oracle Cloud Infrastructure, no public self-serve API is announced, and there are no open prices and latencies. The real way to encounter the technology today is inference through Oracle Cloud Infrastructure and the company's partners. For ML engineers and product teams, practical actions now are tactical: profile your own workloads and understand whether inference is memory-bound, and also keep the inference backend swappable, building an abstraction over providers in the style of OpenAI-compatible APIs so as not to depend on a single supplier. If the statements about LPDDR5X are confirmed, inference will become cheaper, and by the end of 2027 to the beginning of 2028, the normalization of cheap huge context is possible: agents with persistent memory for the entire corpus of company documents, partial abandonment of heavy RAG wrappers, and new classes of products with UX designed for millions of tokens of context.

What is still unknown / limitations

The growth in valuation from just over $1B to $5B in seven months is a financial signal, not a scientific confirmation: the capital market values the narrative and scarcity, not reproducible benchmarks, and the valuation itself says nothing about the truth of the claim of achieving over 90% of available memory bandwidth. The methodology for this measurement is not in the sources: it is not specified which models were used, what batch size, and whether these are dense models or MoE. In addition, high utilization of LPDDR5X bandwidth is not equal to HBM-class absolute bandwidth, since LPDDR5X has a fundamentally lower peak bandwidth per chip. The Asimov chip and Titan system are promises of future hardware, and it is methodologically impossible to evaluate them as current capabilities, and the tape-out may be delayed. Independent public benchmarks for Atlas and a public methodology for measuring memory utilization have not yet been published, so the key technical verification of this story is still ahead.

Sources

Author

Look at AI, editorial team