GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Etched Exits Stealth With $800M, $1B in Contracts, and a Chip Designed to Replace 160 H100s

Etched, a Cupertino-based AI chip startup founded in 2022 by Harvard dropouts Gavin Uberti and Chris Zhu, came out of stealth on June 30 with $800 million raised across four previously undisclosed financing rounds. The latest, a $500 million round that closed in December 2025 at a $5 billion post-money valuation, was led by Stripes. The cap table now includes Jane Street, Hudson River Trading, Jump Trading, Two Sigma, Ribbit Capital, VentureTech Alliance (a fund with a strategic relationship with TSMC), Peter Thiel, Geoffrey Hinton, Andrej Karpathy, Tri Dao, Fei-Fei Li, and Scott Wu.

The company says it has signed more than $1 billion in customer contracts and is actively validating its first rack-scale product with customers.

The Bet

Sohu is purpose-built for exactly one architecture: the transformer. There is no reconfigurability built in for other model types, no general-purpose GPU fallback path. The silicon encodes transformer attention and feedforward computation directly, rather than implementing it on a flexible compute fabric. Etched calls this a deliberate design choice: give up generality, buy throughput and cost efficiency.

The company claims Sohu delivers inference up to 20 times faster than Nvidia’s H100 and says a single Sohu-equipped server can replace up to 160 H100s for transformer workloads. Those numbers have not been independently validated through standardized benchmarks. What is confirmed: Sohu achieved first-pass silicon success on TSMC’s N4P process, it is running models including DeepSeek, Qwen, Mamba, and Llama in customer validation today, and it has more than 400 engineers drawn from Nvidia, Google TPU, Broadcom, SK Hynix, and TSMC.

Two technical claims anchor the architecture story. Low-Voltage Inference runs the compute blocks at less than half the voltage of conventional AI chips, which Etched says allows trillion-parameter sparse MoE workloads to sustain over 80% of peak FLOPs without thermal throttling. Cluster-Scale Memory builds a shared HBM/SRAM hybrid memory pool across the scale-up domain via a proprietary high-bandwidth interconnect, attacking decode latency at the cluster level rather than per-accelerator.

Why the Backers Matter

The cap table is the tell. Jane Street runs some of the most compute-intensive AI inference workloads in quantitative finance and is not a passive investor. VentureTech Alliance’s TSMC partnership gives Etched a strategic alignment with its foundry at a moment when advanced process capacity is constrained. Hinton, Karpathy, and Dao are not writing checks for press coverage — their technical due diligence on transformer-specific silicon is meaningful signal.

The inference chip market has attracted enormous capital in 2026 as the industry’s compute mix shifts from training to serving. Groq was acquired by Nvidia for $20 billion. Google’s TPU 8i is a dedicated inference chip. Cerebras is public at $5.55 billion. Etched is playing the same inference consolidation thesis at the startup tier, but with a harder architectural bet: transformers are the architecture, full stop.

What to Watch

Performance claims will be tested this summer when Sohu ships to production customers. The 20x H100 figure and the 160 H100 replacement claim have no third-party verification yet. If those numbers hold under real workloads, the economics reshape enterprise inference procurement — H100 rentals at CoreWeave and Lambda are the comparison point. If they don’t hold at scale, $1 billion in contracts becomes a headline with asterisks.

The architectural risk is the obvious one. Transformers have dominated every modality for four years. Betting against them in silicon is betting against the next three-to-five years of model development. Etched’s founders made that bet explicitly, and the investors behind them have decided the odds are acceptable.