12 million tokens, linear cost: Subquadratic's bet against the attention tax

The quadratic attention problem has quietly shaped everything you've built with LLMs. RAG pipelines, agentic decomposition, hybrid architectures — these aren't the natural shape of AI systems. They're workarounds. Doubling the context quadruples the compute, so everyone stopped at a million tokens and engineered around the rest. Subquadratic, a Miami-based startup with 11 PhD researchers on staff, launched its first model this week and says it's done with workarounds. Its new architecture — Sub...

Original Source

Read the full article at Dev →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.