DensityAI
Memory-centric inference accelerator for frontier LLMs · DensityAI
DensityAI builds AI accelerators with a novel memory architecture specifically for frontier-scale LLM inference, combining near-memory compute topology of SRAM with HBM density. They target the same inference-only wedge as d-Matrix with explicit focus on beating GPUs on energy and speed for large-model, long-context workloads in datacenter environments.
Why: Nearly identical positioning: inference-only for frontier models, memory-centric architecture, datacenter buyers, targeting speed and energy for large models. Main difference is SRAM+HBM hybrid vs d-Matrix's 3D stacked DRAM approach, but the strategic wedge and buyer profile are almost perfectly aligned.
Oenerga
Memory-native AI infrastructure for frontier model bottlenecks · Oenerga
Oenerga builds memory-native AI infrastructure explicitly designed for frontier model bottlenecks including state movement, KV-cache pressure, and long-context deployment economics. They position as 'post-GPU architecture' targeting the same memory wall problem d-Matrix addresses, with focus on datacenter deployment of large models.
Why: Very similar memory-centric thesis targeting frontier model inference bottlenecks. Explicitly calls out KV-cache and state movement issues that d-Matrix also addresses. Less detail on chiplet disaggregation or software integration strategy, but core technical bet and market timing are closely aligned.
Etched
Frontier inference clusters with co-designed chips and racks · Etched
Etched builds 'frontier inference clusters' by co-designing chips, racks, software, and manufacturing for best-in-class throughput on frontier models. While they share the inference-only wedge and datacenter buyer focus, their emphasis appears more on throughput and full-stack vertical integration rather than d-Matrix's latency-first, modular chiplet approach.
Why: Shares inference-only specialization and hyperscaler buyer focus, but differs significantly in architecture philosophy (vertically integrated clusters vs modular chiplets) and performance dimension (throughput-focused vs latency-first). Less emphasis on memory-centric compute as the core technical bet.
SambaNova
Dataflow architecture for fast decode and agentic inference · SambaNova
SambaNova's RDU (Reconfigurable Dataflow Unit) is purpose-built for fast decode and premium AI inference with their fifth-generation SN50 chip targeting agentic workloads. They emphasize dataflow efficiency and tokens-per-watt, serving datacenter inference at scale, though their reconfigurable dataflow architecture differs from d-Matrix's memory-centric approach.
Why: Strong inference focus with emphasis on decode speed and agentic workloads, targeting datacenter buyers. Dataflow architecture is a different technical bet than memory-centric compute, and less emphasis on chiplet modularity. Performance focus includes both speed and throughput rather than latency-first.
Archality
Memory wall solution for AI inference · Archality
Archality explicitly targets the memory wall problem in AI inference systems, focusing on reducing power and capital waste from constant model weight movement. Very early stage (founded 2025, 3 employees) but shares the core memory-centric thesis, though limited public information on specific architecture or buyer focus.
Why: Shares the memory wall thesis and inference focus, but extremely early stage with minimal public detail on architecture, buyer strategy, or performance dimensions. The core technical insight aligns but execution details are unclear.
Gimlet
Agent-native inference cloud on heterogeneous hardware · Gimlet
Gimlet provides managed inference API for agentic workloads built on heterogeneous hardware, claiming step-change gains in latency and throughput. They target the inference wedge with emphasis on agentic use cases, though they appear to be a cloud service layer rather than a hardware company, making them architecturally different from d-Matrix.
Why: Shares inference specialization and latency focus, with partnership mentioned with d-Matrix for '10x inference speed up.' However, appears to be a managed service/cloud layer rather than hardware company, so technical bet and architecture philosophy differ significantly. Complementary rather than directly comparable.
NeuReality
AI inference optimization with purpose-built NR1 chip · NeuReality
NeuReality's NR1 chip replaces traditional CPU and NIC architectures to optimize AI inference workloads at scale, working with any AI model and accelerator. They target enterprise datacenter inference but position as infrastructure layer that works with existing accelerators rather than replacing them, a different wedge than d-Matrix's standalone accelerator approach.
Why: Inference-focused with datacenter buyers, but positions as infrastructure layer (CPU/NIC replacement) that works alongside accelerators rather than as the primary compute engine. Different wedge and technical bet, though shares the inference scaling crisis timing.