@no_stp_on_snek: https://subq.mildlyconcerning.com

X AI KOLs Timeline News

Summary

This article critically analyzes the claims and timeline of the subQ long-context AI technique, highlighting discrepancies and walkbacks from the original announcement.

@Hesamation https://t.co/5ns8Xe5rhJ
Original Article
View Cached Full Text

Cached at: 05/27/26, 07:20 AM

@Hesamation https://t.co/5ns8Xe5rhJ


Demystifying subQ — A Field Guide

Source: https://subq.mildlyconcerning.com/ Field Notes / v.01Updated: ongoing

Contents

From-scratch frontier model

2026-05-05

2026-05-07

Revised

Demoted to future tense (“we will get there”)

Open-source release of the production model

2026-05-05

2026-05-07

Reversed

“This model won’t be open-source”

52× FlashAttention speedup — scope of headline framing

2026-05-05

2026-05-05

Revised

Prefill-stage only, vs. FA2 on B200s

Base model identity (which open weights were ported)

2026-05-05

2026-05-12

Refused

Question dodged May 7; explicitly refused May 12: “We won’t share many details about the params, base, and training.”

Comprehensive model card

2026-05-05

2026-05-14

Scope reduced + deferred

May 12 X reply: “We won’t share many details about the params, base, and training” (scope reduced). May 14 X reply on the “promised this week” deadline: “Candidly, we may take a little bit longer, as we are working on some cool capabilities we want to be able to showcase.” (timeline deferred).

Third-party validation in model card

2026-05-07

2026-05-11

Released (scope-limited)

Appen Ltd 8-page technical brief delivered. Independence Statement discloses API + auth-key + algorithm-code-review + side-by-side-test access. “Full Report Available on Request” (NDA).

Beta access (early-access list)

2026-05-05

Unreleased

30K+ signups; not yet open as of May 7. “Soon.”

Public release timeline

2026-05-05

Unreleased

May 7 email: “firming up our release timeline.” No specific dates.

Hard selector-efficiency figures vs. DSA

2026-05-05

2026-05-12

Re-promised

Promised “next week” on May 5; not shipped. Re-promised May 12: “We will compare to NSA (DSA) next.”

MRCR 2-needle / 4-needle / non-1M results

2026-05-12

Unreleased

May 12 X reply: “We can follow up with the rest of MRCR. We have more benchmarks coming.”

Perplexity (PPL) figures

2026-05-12

Hedged

May 12 X reply: “I am surprised on the request for perplexity, given how little it tells you and the dataset bias. We could maybe follow up with that.”

12M-token benchmarks (the headline context-window claim)

2026-05-05

2026-05-12

Re-promised

Launch headline was 12M context. Appen brief stops at 1M. May 12 X reply to “why benchmark at 1M, I thought the whole point was 12M?”: “12MM-token benchmarks coming next!”

SubQ-branded models (own model releases)

2026-05-12

Promised

Asked “does this work with current models? Basically frontier labs can use subq for inference, right?” Whedon May 12: “We are going to release our own models!”

Continuous evaluation via LayerLens Stratix

2026-05-14

Promised

Whedon X 2026-05-14: “Results and future evaluations will be published publicly.” Per subq.ai post: “Results coming soon at stratix.layerlens.ai.” No specific date.

Decode-stage speedups

2026-05-05

Unreleased

Described as shipping later

Technical paper / arXiv submission

2026-05-05

Unreleased

Not findable on arXiv as of May 7, 2026

Public training code & kernels

2026-05-05

2026-05-07

Reversed

“We can only go so far”; future open-sourcing hedged

Initial claim at launch

Walk-back, modification, or reversal

2026-05-05

Claim · from-scratch frontier model

“Introducing subQ” — first fully sub-quadratic frontier

Launch positions SubQ as the “first fully sub-quadratic frontier” model. The framing reads as a frontier-class model trained from scratch with a novel architecture.

2026-05-07

Modification · revised to future tense

“Correct. We will get there on the from-scratch!”

Replying to a summary describing SubQ as a DSA-variant retrofit on open-weight base with CPT/SFT/RL on top, Whedon answers:*“Correct. We will get there on the from-scratch!”*The shipped production model is confirmed not from-scratch.

2026-05-05

Claim · open-source plans

“Open-source plans” for the first SubQ model

Launch-day reply describes what is planned for public release — weights, training code, kernels — and the rough sequencing.

2026-05-07

Modification · reversed

“This model won’t be open-source”

Reply to Greg Horvay:“Unfortunately, this model won’t be open-source, but we do intend to contribute to the open-source community.”Follow-up on divulging the improvements:“We are getting pressure on this, and we can only go so far. However, we have better things coming, so we can maybe open-source the old things over time, or we can open-source some of our lessons/data/tools.”

2026-05-05

Claim · 52× FlashAttention speedup

Headline: “52× faster than FlashAttention”

Launch headline frames a 52× speedup vs. FlashAttention as a top-line performance number, without scope qualification in the announcement post.

2026-05-05

Modification · same-day clarification

52× applies to prefill only, vs. FA2; decode speedups shipped later

Replying to Martin Shkreli, Whedon clarifies the 52× number is a prefill-stage speedup measured against FlashAttention 2 on B200s. Decode-stage speedups described as shipping in a later release.

2026-05-05

Claim · novel architecture, frontier model

Frontier model, “Subquadratic Sparse Attention” novel architecture

Launch positioning. The launch and blog do not name a base model and the framing reads as architecture-defined frontier model.

2026-05-05

Modification · ported open-source weights

Production model is ported open-source weights + CPT/SFT/RL

Whedon clarifies on May 5:*“Yes, that is not correct. We do CPT, SFT, and RL after porting weights over to a new architecture.”Subsequently on May 7, asked specifically which open base, Whedon replies:“I have a few ideas about how to do this in the model card. I will make sure there is something side-by-side!”*Base identity remains undisclosed.

2026-05-14

Claim · third-party continuous evaluation

SubQ × LayerLens — continuous evaluation partnership announced

Whedon X (verbatim):*“We’ve partnered with @layerlens_ai to continuously evaluate SubQ across nearly 100 benchmarks and 200+ frontier models on Stratix... Results and future evaluations will be published publicly.”*subq.ai blog post frames the engagement as an “independent evaluation layer.” No publication date stated; no commercial terms or prior-relationship disclosure in the announcement.

Alex Whedon’s main SubQ announcement post on XMain launch · video

The announcement post

The original post that kicked everything off, with the launch video. Conversation ID anchor — most replies threaded here.

x.com/alex_whedon/status/2051663268…↗

SubQ early access announcement postAccess · agent

Early access & agent link

How to actually get hands-on with the product — the early-access announcement plus the agent endpoint.

x.com/alex_whedon/status/2051663274…↗

Technical blog post announcement on XMost important · technical

Technical blog post announcement

The single most useful post for understanding the product. Linked and reposted by Alex in multiple threads — duplicates included below for context.

Reply to Martin Shkreli — 52x is for prefill speedReply to Martin Shkreli

52× prefill / decoding clarification

Pinning down what the speedup number actually refers to — prefill, decoding, or both — in response to skeptical questioning.

status/2051686740…↗

Reply to Vincent — model card coming next weekArchitecture details

Model card & blog details

Pointer to the model card, plus inline detail on how SSA is described in the technical write-up.

status/2051702531…↗

Reply to Arthur — dynamically select token relationshipsReply to Arthur · long-range deps

Recall & selector mechanism

How the selector chooses which tokens to attend to and how recall is preserved across long contexts — answer to a long-range dependency question.

status/2051719125…↗

Reply to Faruk Guney — Subquadratic Sparse Attention is a unique variant we createdPositioning

SSA avoids the usual tradeoffs

The pitch: SSA gets long-context efficiency without paying the quality tax that other sparse / linear methods do.

status/2051700602…↗

Reply to toriset — a linear form of sparse attention, similar to DeepSeek but without the quadratic bottleneckClarification

Novel sparse attention — not linear attention

Drawing the line: SSA is sparse attention with novel selection, not a linear-attention variant. Important for anyone benchmarking mentally against Mamba / RWKV / etc.

status/2051719585…↗

Whedon: We have done FA3. FA3 doesn’t provide a speedup relative to FA2 on B200s. FA4 is in sight. Plus follow-up: Yeah, FA3 is built for Hopper chips specifically, so you don’t see the speedup on Blackwells.Reply · FA3 not used as baseline

FA3 does not provide a speedup over FA2 on Blackwell B200s — CTO statement

Whedon X 2026-05-12 (verbatim): “We have done FA3. FA3 doesn’t provide a speedup relative to FA2 on B200s. FA4 is in sight, and we are looking at what gains we can capture from studying FA4 too.”

Same-day follow-up (verbatim): “Yeah, FA3 is built for Hopper chips specifically, so you don’t see the speedup on Blackwells!”

Independent verification: Dao-AILab/flash-attention issue #1853 documents that FA3 errors out on Blackwell (sm_100): “FA3 is only supported on devices with compute capability >= 8 excluding 8.6 and 8.9 and Blackwell archs (>=10).” Confirms the FA2 baseline choice in the Appen brief.

sdmat asks if SubQ is really linear or N·K; Whedon replies: N*K time, and unlike SSMs, there is still explicit retrievalReply to sdmat · complexity admission

N×K time, explicit retrieval — the internal-RAG admission

sdmat asked the right question: if SubQ is strictly linear, the router is O(1) per query in N and logically must have a capacity bound similar to SSMs — just moved from state to selection. So is “linear” technically N polylog N, or N·K sized for this regime? Whedon answered: “N*K time, and unlike SSMs, there is still explicit retrieval.” Two admissions in one sentence. First, the complexity is N×K, not strictly linear in N — sdmat’s framing was correct. Second, the architecture has an explicit retrieval/selection step over the context. Different storage mechanism than SSMs (selected tokens vs compressed state), same fundamental capacity ceiling. The “novel architecture” framing survives only if learned-K-token-selection is treated as categorically different from learned-K-dim-state-compression. From a capacity-theoretic view, it isn’t — it’s internalized RAG.

status/2052170998…↗

Reply to nanou — when the constraint goes away, what people build is differentPipeline impact

Long-context pipelines change

The argument for why entire workflows — not just single inferences — get rebuilt when long-context becomes cheap enough to be the default.

status/2051668090…↗

Long-context demand is about to go higher; the premium charged on top of it is going downPricing thesis

Long-context demand & pricing premium

The pricing forecast in two sentences: long-context demand should rise as workflows shift to assume cheap context, and the premium charged on top of it compresses.

status/2051674857…↗

Reply to Poonam Soni — the unit economics of agentic products were quietly upside down. Well, not anymoreReply to Poonam Soni

Agentic product unit economics

On the implication for agent products: 5% of frontier price at 12M tokens flips assumptions a lot of agentic stacks were quietly built on. Alex’s own framing — “quietly upside down. Well, not anymore.”

status/2051668955…↗

Reply to Rob Parker on May 7 — confirms the model is intended as a plug-and-replace agentic LLM substituteReply · agentic plug-and-replace

Plug-and-replace agentic positioning

Asked directly whether the SubQ model has agentic abilities usable as a plug-and-replace for current LLMs, Whedon answers with a single word:*“Yes.”*A confident commitment to drop-in agentic deployment, made before any model card, weights, or agent-specific eval beyond the launch-table SWE-Bench number have been published.

status/2052454133…↗

Reply to Linus Ekenstam — workloads that historically driftQuality regime

Drift & degradation workloads

The class of long-running / long-context tasks where standard attention quality starts to drift — and where Alex argues SSA holds up.

status/2051673396…↗

Reply to Jas Oberoi — comparing against FlashAttention sets the bar higher, not lowerReply to Jas Oberoi

FlashAttention as the higher bar

On why benchmark against FA at all: sparse attention reduces FLOPs in theory, but the implemented speedup is often lower — and sometimes negative. Comparing against FA forces the comparison to be about real, measured performance.

status/2051861776…↗

Reply to Pratyush Tiwari — FA4 wasn’t out when the benchmarking was done; integration in progressReply to Pratyush Tiwari

FA4 integration in progress

On why the comparison used FlashAttention-2 on B200s instead of something newer: FA4 wasn’t out yet when the benchmarks were run. FA4 improvements are being integrated into the inference stack now.

status/2051771504…↗

SubQ’s official benchmark comparison table — SWE-Bench, RULER, MRCR v2 vs Gemini 3.1 Pro, Opus 4.6, Opus 4.7, GPT-5.4, GPT-5.5Company source · published comparison

SubQ’s official benchmark table

The comparison table on subq.ai. SubQ 1M-Preview vs. Gemini 3.1 Pro, Opus 4.6 / 4.7, GPT-5.4 / 5.5 across SWE-Bench Verified, RULER@128K, and MRCR v2. Notable: Opus 4.6 outscores SubQ on MRCR v2 (78.3 vs. 65.9) — a comparator that does not appear in the launch communications.

subq.ai↗

Reply to Stella Biderman — Light DSA, dynamically selected token relationships, but without the high expense of the lightning indexerReply to Stella Biderman

DSA — same idea, lower cost

On how the mechanism works: dynamically selected token relationships, like Light DSA — but without the high expense of the lightning indexer that makes DeepSeek’s variant quadratic.

status/2051784892…↗

Reply to Andy / PositivFuturist — DeepSeek Sparse Attention is helpful for explaining what we didReply to PositivFuturist

DSA as the teaching analogy

DSA is the easiest reference point Alex has found for communicating the mechanism: same dynamic selection of token relationships, but without the quadratically-scaling lightning indexer.

status/2051767936…↗

Marketing framing

Marketing vs. the DeepSeek paper

Response to the framing that the launch is mostly marketing built on top of DeepSeek’s paper — Alex’s pushback on what’s genuinely new.

status/2051834849…↗

Third-party · SWE-Bench comparison

SWE-Bench: SubQ 81.8 vs. Opus 4.7 87.6

Per Jake Cuth’s verification analysis, on the SWE-Bench benchmark itself:“SubQ’s 81.8 SWE-Bench is below Opus 4.7’s 87.6…The ‘outperforms’ framing is technically true on a cherry-picked-per-model basis, not on a same-benchmark basis.”

jakecuth.com (analysis)↗

Reply confirming base model: open-source weights ported into the SSA architecture, then CPT, SFT, and RL appliedReply · base model confirmation

Open-source weights, ported and post-trained

The clearest statement on how the production model was built: weights from open-source models as a starting point, ported into the SSA architecture, then continued pretraining, supervised fine-tuning, and reinforcement learning. Framed as a function of funding and company maturity. From-scratch experiments, by Whedon’s own description, are confined to smaller scale.

status/2051755590…↗

Whedon reply on May 7 confirming SubQ is not a from-scratch frontier model: ‘Correct. We will get there on the from-scratch!’Follow-up · from-scratch admission

“We will get there on the from-scratch”

Two days after launch, Whedon explicitly confirms what the base-model port already implied: the production SubQ model isnota from-scratch frontier model. Replying to a summary that calls it a DSA-variant retrofit on someone else’s base with CPT + SFT + RL on top, Whedon answers*“Correct. We will get there on the from-scratch!”*— placing from-scratch in the future tense. Anchors the $29M seed (~10% of a frontier pretrain) against what was actually shipped on May 5: a clever retrofit, not a ground-up frontier system.

status/2052452620…↗

Reply to will depue — We are working on the memory problem too!Memory · kernels

Memory & infra work

Two posts on the systems work behind SSA — memory layout, kernel choices, and what made the long-context numbers possible.

Reply to will depue — linear complexity for compute and memory, but training and inference infra still needed work to serve >1MM tokensReply · willdepue thread

willdepue thread — complexity, training, porting

The most-pressed thread on the architecture: complexity class (linear in compute & memory), the from-scratch vs. ported question, and why infra work was the real bottleneck even with sub-quadratic theory.

Third-party · GPU contract terms

Digi Power X $19.6M B300 contract — effective May 15, 2026

Per Jake Cuth’s verification analysis:*“$19.6M GPU contract Twenty-four-month bare-metal Blackwell B300 rental from Digi Power X (NASDAQ: DGXX). Signed April 20, 2026. Effective May 15, 2026. 15% upfront.”Cuth additionally notes:“Stock-moving infra reveal — The Digi Power X contract moved DGXX +19% on announcement. The disclosure benefits both counterparties.”*The dedicated B300 capacity does not begin until May 15, 2026 — ten days after the May 5 launch and consistent with Whedon’s own May 7 statement that the launch-time compute is “neocloud allocations we hobbled together” on ~64×B200 steady-state.

jakecuth.com (analysis)↗

Reply · compute scaling cadence

“We just doubled our compute today. That isn’t saying that much.”

Asked by**@OliwierMako**, May 7, 2026:“Do you plan to get more compute to potentially compete with Anthropic or OpenAI or do you plan to sell this breakthough?”

Whedon’s reply, May 7, 2026:“We are acquiring more compute. We just doubled our compute today! That isn’t saying that much. We will be at triple where we were yesterday early next week. We have larger ambitions to throw compute at.”

On the record: yesterday’s footprint × 2 = today’s; yesterday’s footprint × 3 by next week. Mapped against Whedon’s earlier-today disclosure of ~64×B200 steady-state on hobbled neocloud, this scales to roughly 128×B200 today, 192×B200 by next week, with the $19.6M Digi Power X B300 contract effective May 15.

status/2052480135…↗

Whedon May 7 reply — neocloud allocations hobbled together, 448xB200 cluster for 10 days then pre-empted by OpenAI, ~64xB200 most of the time, helpers Cudo Hydra LambdaLabs GMIReply · compute footprint

Compute scale: hobbled neocloud, 64×B200 steady-state

Whedon volunteers the actual training-compute story:*“We are using some neocloud allocations we hobbled together. We did have a 448xB200 cluster for 10 days and thought our problems were solved, and then we got pre-empted by OpenAI. Most of the time, we are working on ~64xB200 clusters. Key helpers have been Cudo, Hydra, and LambdaLabs.”*A second post adds GMI. This is the compute receipt that converges with the from-scratch admission via independent evidence: a steady-state of ~64 B200s is a CPT / SFT / RL footprint, not a frontier-pretrain footprint, which sits on thousands of accelerators for weeks-to-months. The 448-B200 burst was real frontier-class compute, but only for 10 days before OpenAI bumped them off the same neocloud — being pre-emptable means low priority on the shared pool. Together: matches the “ported open weights + post-train” architecture story exactly.

Reply in eliebakouch thread — with additional funding, intend to train a model without any form of dense attention; leading candidates not attention-based at allReply · eliebakouch thread

Future training: “not attention-based at all”

The forward-looking commitment: with additional funding, train a model without any form of dense attention. Whedon notes the leading candidates for that future model “are not attention-based at all in the strictest sense” — a stronger claim than the launch architecture itself makes.

status/2051756365…↗

Reply — research on memory validated at small scale, requires training from nothing; selector is far more efficient than DSA, hard figures next weekReply · memory + selector

Memory research: “train from nothing”

Two future-direction notes in one reply: ongoing memory research validated at small scale, described as “very radical” and requiring training from nothing; and the selector efficiency claim vs. DSA — far more efficient, with hard figures deferred to the model card next week.

status/2051764258…↗

Reply about open-source plans for the first SubQ modelOpen source

Open-source plans

What’s planned for release publicly — weights, training code, kernels — and the rough sequencing.

status/2051702240…↗

Reply · technical paper timeline

Technical paper: “Next week”, partial details only

@paperbyhand, May 7, 2026:“when will your research paper be dropping?”

Whedon’s reply, May 7, 2026:“Next week! We pulled some of the content into the technical blog post as a teaser. We will not be able to share all the details of how it works but can show some more analysis and ablation.”

On the record: the existing technical blog is described by the CTO as a “teaser”, the forthcoming paper is scheduled for “next week”, and Whedon states it will not share all details of how the architecture works. Follow-up from**@pastimpress**in the same thread asks for the researchers it will be attributed to and the publishing organization; not yet answered.

status/2052480123…↗

Corporate email · early-access list

May 7 email to beta signups: “Claims mean nothing without proof”

Email sent May 7, 2026 at 4:12 PM MT from**[email protected]**to early-access signups, signed “Alex & The SubQ Team” (CTO; no CEO signature). Verbatim:

“The response over the last 48 hours have been massive — and we’ve seen the excitement and the skepticism. Both are fair. We’d rather earn trust than ask for it.

Next week we’ll publish our model card with additional data and third-party validation. Claims mean nothing without proof, and we plan to back ours up.

We’re also firming up our release timeline and plan to start opening beta access soon. Hang tight. It’ll be worth the wait.”

Three new commitments on the record from corporate distribution: model card next week withthird-party validation(a new specific commitment beyond the prior “model card coming soon” framing); release timeline still being firmed up; beta access not yet open. The phrase “claims mean nothing without proof, and we plan to back ours up” is an implicit corporate-channel acknowledgment that current claims lack independent proof.

Reply · competitive positioning

“Absolutely” — plans to compete with Anthropic / OpenAI

@OliwierMako, May 7, 2026:“That sounds really good. But I got one more question. Do you plan to compete with Anthropic/openAI?”

Whedon’s reply, May 7, 2026:“Absolutely”

On the record: public commitment to compete with hyperscaler-tier frontier labs.

status/2052487466…↗

Reply · roadmap

“The mission is far from complete with sparse attention”

@ai_hyperbull, May 7, 2026:“Do you think getting Aquihired by a lab is a logical conclusion to your work?”

Whedon’s reply, May 7, 2026:“We are in this for the mission, and the mission is far from complete with sparse attention.”

On the record: the current sparse-attention architecture is described by the CTO as not the end-state of the mission. Pairs with the May 7 “we will get there on the from-scratch” statement — the shipped product is framed as intermediate.

status/2052478738…↗

Whedon May 7 reply — Unfortunately, this model won’t be open-source, but we do intend to contribute to the open-source communityFollow-up · open-source clarified

“This model won’t be open-source”

The hard clarification on top of the earlier “open-source plans” framing. Asked whether tools like LM Studio and Ollama would need updates to support SubQ, Whedon answers:*“Unfortunately, this model won’t be open-source, but we do intend to contribute to the open-source community.”A follow-up, asked whether the improvements themselves would be divulged so other open-source models could employ them, draws an explicit moat-protection answer:“We are getting pressure on this, and we can only go so far. However, we have better things coming, so we can maybe open-source the old things over time, or we can open-source some of our lessons/data/tools.”*Future open-sourcing is hedged (“maybe”) and conditional on the team having something newer to keep proprietary.

status/2052453884…↗

Roadmap

Future pipeline / vision

Where SSA-style models go from here, and what the longer-term product surface looks like beyond the current launch.

status/2051783974…↗

Repo · agent

Repo & agent handling

Alex has been replying to a steady stream of “drop the repo” asks throughout the main thread. There isn’t one canonical post — sort his profile byLatestand search for"Drop the repo"to scan the most recent answers.

x.com/alex_whedon↗

Technical blog · canonical read

How SSA makes long-context practical

The full technical write-up. Linked from most of Alex’s posts. If you read one thing, read this.

subq.ai/how-ssa-makes-long-context-practical↗

Profile · live updates

Alex Whedon on X

For anything new or missed: visit the profile directly and sort byLatest. The skeptic threads and deep-dives keep generating useful replies.

x.com/alex_whedon↗

Third-party · founder backgrounds

CEO and CTO professional histories

Per Jake Cuth’s verification analysis:

CEO Justin Dangel:“Founded Voter.com (1998), Goji (2008), co-founded Firefly Health and Ready Responders (GV / Founders Fund backed).”On research credentials:“No ML-research CEO — Justin Dangel’s record is healthcare, insurance, consumer. No publication record, no prior AI work.”

CTO Alex Whedon:“Ex-Meta software engineer. Former Head of Generative AI at TribeAI. NeurIPS / AAAI / ICLR 2026 attendee.”

jakecuth.com (analysis)↗

Independent · The New Stack

“The context window has been shattered”

Frederic Lardinois at The New Stack — the most thorough launch piece. Includes the launch-interview benchmark numbers (RULER 97.1, MRCR v2 83, NIAH 12M 92.1) which, on closer inspection, differ materially from the production numbers in SubQ’s own published table. Also surfaces the “harness as much as model” SWE-Bench caveat from SubQ’s own technical paper.

thenewstack.io↗

Independent · VentureBeat

Researchers demand independent proof

VentureBeat’s launch piece foregrounds the skeptical reception. Confirms Whedon’s open-source-weights admission, walks through the cherry-picked benchmark structure, and draws the parallel to Magic.dev’s August 2024 trajectory: 100M-token claim, ~$500M raise, no public deployment by early 2026.

venturebeat.com↗

Independent · SiliconANGLE

Launches with $29M, founder interview

Kyt Dotson at SiliconANGLE — the company-friendly launch piece built around a Dangel + Whedon interview. Useful as the preferred narrative: $29M raise at a reported $500M valuation, three products (API, SubQ Code, search), running on neoclouds rather than the major hyperscalers.

siliconangle.com↗

Skeptical analysis · pre-launch

Debunking claims about subquadratic attention

Vladimir Ivanov, LessWrong, January 1, 2026 — four months before the SubQ launch. The foundational skeptical framework for evaluating subquadratic claims. Argues the genre fails one of two ways: quadratic in implementation (Kimi Linear, DSA still rely on a quadratic indexer) or underperforming at frontier scale (pure Mamba / RWKV). Identifies DSA’s lightning indexer as the specific quadratic bottleneck — the one Whedon now claims SSA removes.

lesswrong.com↗

Reply · investor public defense

Investor Justin Mateen: “It would be foolish to share every detail”

@willdepue(OpenAI), May 5, 2026, after reading SubQ’s technical materials:“Let’s read the technical report. TLDR; No real answers on how their method works. Doesn’t make me feel better about it.”

@justinmateen(Tinder co-founder; named SubQ investor in the Refresh Miami launch coverage), May 5, 2026:“I really respect your take and thoughtfulness. It would be foolish to share every detail of how the sparse attention architecture works publicly. Full disclosure, I’m a large shareholder.”

On the record: a named SubQ investor publicly defends non-disclosure of architectural detail, with explicit shareholder-interest disclosure. The substantive defense of the closed-architecture stance comes from the cap table rather than the founders. Aligns with Whedon’s later (May 7) “we can only go so far” framing on open-sourcing improvements.

status/2051809058…↗

Reply · CTO endorses third-party critique

“We shared this internally on Slack earlier this week”

@ItsCuthulhu(Jake Cuth, the analysis author), May 7, 2026, sharing his own piece:“Sidenote: People keep calling it your company a scam but I made an interactive dashboard and do not see evidence of this. The issue lies in people not having enough verification yet, rather than being a scam. Just so you’re aware.”

Whedon’s reply, May 7, 2026:“We shared this internally on Slack earlier this week! We all thought this was a great breakdown. Thanks for the post! Yeah, we just need a little more time for product release and short-context benchmarks.”

On the record: SubQ’s internal team had Cuth’s verification analysis — the one flagging the SEC-vs-press valuation gap, the 17-point MRCR v2 production-vs-research delta, the SWE-Bench “outperforms” framing relative to Opus 4.7’s 87.6, the unverified 12M context and 1,000× compute claims, and the no-arXiv / no-leaderboard footprint — circulating internally on Slack before May 7. The CTO endorses it as a “great breakdown” in public and frames the open items as time-to-deliver, not factual disputes.

Third-party analysis · post-launch

“Real company, overhyped claims, reserve judgment”

Jake Cuth’s independent post-launch analysis, framed explicitly asnot a hit piece. Categorizes SubQ as operationally legitimate but performance claims unverified.**Verified facts:**Miami startup, SEC Form D (Feb 2026), 29M seed at ~500M valuation, $19.6M GPU contract with NASDAQ-listed Digi Power X, and one unnamed third party confirming production scores RULER 128K = 95.0 and MRCR v2 = 65.9.**Unverified:**the 12M context window, the ~1,000× compute reduction, the “fully sub-quadratic” O(n) attention claim, and a self-reported research-mode MRCR v2 of 83.0 — nine points above frontier models. Cuth flags the 17-point internal gap between SubQ’s production (65.9) and research (83.0) numbers as “itself larger than most frontier model differences,” warranting independent reproduction before credibility. Cites the Magic.dev precedent — similar claims in 2024, no validation 21 months later — as the cautionary case.

jakecuth.com/work/subquadratic-lab↗

Archive snapshot

Technical blog (archived)

Archive.is snapshot of the canonical technical blog post, “How SSA makes long-context practical.” Useful as a stable reference if the live page is updated, paywalled, or removed.

archive.is/pgxaW↗

Kamradt May 14 reply: ‘arc-agi specific verified scores. They worked with Appen for the report mentioned in the OP tweet’Reply · Kamradt scope clarification

Kamradt clarifies: ARC-AGI-specific verified scores; Appen brief is “the report” SubQ has

Asked whether the “verified scores in a few weeks” mentioned in his earlier meeting report were ARC-AGI-specific or first verified third-party scores in general, Kamradt X 2026-05-14 11:25 AM (verbatim):

“arc-agi specific verified scores. They worked with Appen for the report mentioned in the OP tweet”

What this clarifies: · The “few weeks” verification work is ARC-AGI-specific, not a broader verification engagement · Kamradt characterizes the existing Appen brief as “the report” SubQ has — SubQ’s existing third-party validation in Kamradt’s framing · Kamradt is not endorsing the Appen brief’s independence framing; he is reporting that it exists

Kamradt clarification reply↗

Whedon May 14 LayerLens partnership announcement (the repost; 10:12 AM with frontier typo corrected from the deleted original posted earlier the same day)Partnership announcement · LayerLens

SubQ × LayerLens — “continuous evaluation” partnership announcement

Announced 2026-05-14. The original tweet (status/2054929047692399063) was deleted and reposted at 10:12 AM the same day as status/2054942865390780670, with the typo “froniter” corrected to “frontier.” The original is no longer publicly accessible.

Whedon X (verbatim):“We’ve partnered with @layerlens_ai to continuously evaluate SubQ across nearly 100 benchmarks and 200+ frontier models on Stratix. The goal is to continuously improve performance while reinforcing a shared commitment to transparency, auditability, and responsible model assessment. Results and future evaluations will be published publicly. Link below to learn more.”

subq.ai blog post (verbatim excerpts): · Framing: an “independent evaluation layer” vs internal assessments · Eval scope: “retrieval accuracy at depth, positional consistency across varying context lengths, and synthesis from extended inputs” + “reasoning, coding, instruction following, and tool use evaluations” · All future releases will use “Stratix Enterprise going forward” · Will publish “methodology, findings, strengths, and limitations” including “the parts that show limitations” · “Results coming soon at stratix.layerlens.ai” — no specific date

About LayerLens (verifiable from layerlens.ai public materials): · Vendor-neutral evaluation platform headquartered in Miami, FL · Founded 2024, $6.5M pre-seed disclosed (Collab+Currency, Generative Ventures, Nazare, Velocity Capital, Web3.com Ventures + 1 undisclosed) · Public model + benchmark counts vary across their own marketing pages: 160–200+ models, 52–100+ benchmarks · Stratix product offered with free + paid (Stratix Premium) tiers · CEO Archie Chaudhury, President Jesus Rodriguez

Not stated in the announcement (verifiable absences): · Specific benchmark names targeted for SubQ · Publication schedule · Whether 12M-context evaluation is in scope · Whether SubQ provides API-only access or also algorithm-code-review / side-by-side-test access analogous to the Appen brief · Whether SubQ pays LayerLens, gets the engagement free, or has another arrangement · Whether the model identity / params / base will be disclosed to LayerLens or remain undisclosed (per Whedon’s 2026-05-12 refusal) · Whether any prior commercial, advisor, investor, or employment relationship existed between SubQ and LayerLens entities or their principals before this engagement

Whedon May 14 reply on LayerLens methodology: ‘They actually maintain all of that themselves. It is a pretty cool platform. Everything is evaluated in a completely flat and transparent way. For example, for SWE-Bench Verified, they use Mini-SWE-Agent for all of the models. This means that all of the models score lower than advertised. You can actually run your own model comparisons there independently.’Reply · LayerLens methodology stance

Whedon on LayerLens methodology — “completely flat and transparent” (May 14)

Asked whether SubQ would supply custom implementation code, scoring functions, or eval traces to LayerLens for any of the ~100 benchmarks (parallel to the Appen brief’s disclosed access conditions), Whedon X 2026-05-14 (verbatim):

“They actually maintain all of that themselves. It is a pretty cool platform. Everything is evaluated in a completely flat and transparent way. For example, for SWE-Bench Verified, they use Mini-SWE-Agent for all of the models. This means that all of the models score lower than advertised 😄. You can actually run your own model comparisons there independently.”

What this commits SubQ to publicly: · LayerLens controls all eval methodology, no SubQ code supplied · Standardized harness across models (Mini-SWE-Agent for SWE-Bench) · Models score lower than advertised under this methodology — including, presumably, SubQ · Independent reproducibility claim: “you can run your own model comparisons there independently”

**For the reader to compare against:**the adjacent “LayerLens prior continuous evaluation precedent” card documents that the publicly-posted CoreThink-Reports PDF on LayerLens’s own GitHub describes a different arrangement for the headline ARC-AGI-2 benchmark (vendor-supplied custom implementation accepted) and for BFCL v3 / BIRD-CRITIC (vendor traces accepted as validation). The two LayerLens engagement profiles will be visible side-by-side once the SubQ results land.

Whedon May 14 reply↗

LayerLens prior “continuous evaluation” precedent

CoreThink Reports — public PDF on LayerLens’s own GitHub org

Source:github\.com/LayerLens/CoreThink\-Reports(created 2025-10-13, publicly accessible).

Repo contains one PDF: “CoreThink SRM Models Report.docx (3) (2).pdf” — the only published precedent of a LayerLens vendor “continuous evaluation” arrangement found in public materials as of 2026-05-14.

Verbatim from the PDF: · ARC-AGI-2 (headline 23.9%, ranked #1):“CoreThink provided a custom implementation with their reasoning layer and verifier functions”that LayerLens“reviewed and executed.” · BFCL v3 (vendor reported 58.5%, LayerLens reported 56.5%):“We were able to validate these figures by reviewing the provided traces and confirmed they fall within the accepted error bounds.” · BIRD-CRITIC (vendor reported 37.2%, LayerLens reported 32.38%): same footnote pattern ·“We will soon be adding CoreThink’s models to our public leaderboard”

Verifiable absences (as of 2026-05-14): · CoreThink models not present on stratix.layerlens.ai/models public leaderboard (~7 months after the promise) · CoreThink’s claimed 23.9% ARC-AGI-2 score not present on the official ARC Prize leaderboard · PDF contains no commercial-terms disclosure

Paid attestation · scope-limited

Appen — “Subquadratic – Model Performance & Architecture Evaluation”

8-page technical brief published 2026-05-11. Hardware: NVIDIA B200, CUDA 13.0, PyTorch 2.11.0, bfloat16. Mean of 5 runs, 3 warmups.

Reported numbers (verbatim from the brief): · Wall-clock @ 1M tokens: 21,410.51 ms FA2 → 380.96 ms SSA (56.2× end-to-end prefill speedup vs FA2) · FLOPs @ 1M tokens: 9,095 TFLOPs FA2 → 144.9 TFLOPs SSA (62.8× reduction; validated via torch.profiler within 0.7–3.9%) · RULER 95.6% (13-task average; brief headlines this under the heading “QA & Word Extraction”) · MRCR 8-needle @ 1M: 86.2% — brief text: “the model either retrieves all eight needles correctly or misses entirely — a bimodal pattern typical of sparse-attention architectures.” · SWE-Bench Verified: 81.8% with extended thinking enabled

Independence Statement (verbatim, page 2):“For benchmarks, access was scoped to Subquadratic’s API endpoints and authentication keys only; no model weights, training data, fine-tuning configurations, or benchmark ground-truth labels were provided in advance.For wall clock and FLOPs analyses, Appen got access to the key algorithm code, did a code review, and was able to run side-by-side tests.

Pre-existing relationship disclosure(Whedon X 2026-05-12, defending Appen’s reputation):*“I have been working with Appen since 2019. They have done dataset creation and benchmarking for many frontier labs and are a publicly traded company. They have 1MM followers on LinkedIn.”*7-year prior client relationship between SubQ’s CTO and the attesting party.

**Not contained in the 8-page brief:**PPL, decode tokens/sec, KV memory in GB, model identity, parameter count, base model, training data, comparisons to any baseline other than dense FA2 (no NSA, MoBA, DSA, MInference, RetrievalAttention, Quest, H2O, StreamingLLM, SnapKV, KIVI, FA3, FA4). Brief notes “Full Report Available on Request” (NDA).

Discrepancy · internal contradiction

What does SSA stand for?

The blog calls it “Subquadratic Sparse Attention” in the lede, then “Subquadratic Selective Attention” in the architecture section — same post, no footnote reconciling the two. Launch tweets and the website hero use Sparse. The technical blog uses both. The community has no canonical name to cite.

Discrepancy · training methodology

“Pre-training” or continued pretraining on open-source weights?

The blog describes a three-stage training process beginning with “Pre-training [that] establishes base language modeling capability.” Reads as from-scratch. On X, when pressed by the willdepue / eliebakouch threads, Whedon clarified the model is initialized from open-source weights, ported into the architecture, with CPT/SFT/RL applied — not pre-trained from scratch. The X clarification was forced by community pressure; the blog’s framing is undisturbed.

Discrepancy · marketing vs data

Does SubQ “keep up with frontier” on MRCR v2?

The blog claims SubQ “keeps up with frontier dense-attention models on MRCR v2.” Lower in the same post, the blog’s own table shows SubQ at 65.9%, Opus 4.6 at 78.3% (12.4-point gap), and GPT-5.5 at 74.0% (8.1-point gap). The marketing claim and the data table contradict each other within a single document.

Per the blog (both lede and table)↗

Discrepancy · framing

“Well in the range of Opus 4.6 at 78”

The MRCR v2 results section frames SubQ’s 65.9% as “well in the range of Opus 4.6 at 78.” The deficit is 12.4 points. By comparison, GPT-5.5 (74.0) and Opus 4.7 (32.2) appear in the same table but are excluded from the “in the range” framing. The blog acknowledges the comparators that flatter SubQ and elides the ones that don’t.

Per the blog↗

Discrepancy · architectural claim

“Does not approximate attention” vs. the trained selector

The blog states SSA “does not approximate attention. It restricts attention to the positions that actually carry signal, and skips the rest.” This frames the mechanism as deterministic exact attention. On X, Whedon clarifies the selector is trained against retrieval problems with distractor documents — meaning it’s a learned component with no formal worst-case guarantee. The blog never addresses what happens when the selector misses a critical position.

Discrepancy · benchmark numbers

The 17-point gap between blog and launch coverage

The blog’s MRCR v2 production score is 65.9%. The number quoted in The New Stack’s launch coverage is 83 — “beats OpenAI by nine points.” Same model name, same benchmark, different numbers, no version disclosure in either venue. RULER 128K shows a smaller gap (97.1 vs. 95.0). The 12M NIAH score of 92.1 appears in TNS coverage and not in the blog at all.

Discrepancy · deferral pattern

“Coming soon” for everything load-bearing

The blog’s own header note opens with: “A comprehensive model card is coming soon!” What’s been deferred to that document: hard figures on selector efficiency vs. DSA, broader benchmarks beyond the three reported, 12M-context evaluation, decoding speedups, and any architectural detail that would let an outside researcher reproduce the claim. The technical blog is itself a teaser for the technical document — and the technical document does not exist yet.

Per the blog header note↗

Third-party · SEC Form D

$500M valuation is single-source, not in the SEC filing

Per Jake Cuth’s verification analysis:*“SEC Form D filed February 2026. Public record via StreetInsider. Confirms the offering exists.”On the valuation specifically:“The ’500M valuation' comes from one outlet \(The New Stack\) and is not corroborated by the SEC Form D, which reports the size of the offering but not the valuation directly\.”*The widely\-quoted “29M at $500M” framing relies on a single press source for the valuation half of that sentence.

Discrepancy · launch graph disproportionality

CTO acknowledges launch graph was disproportional

Critic**@bomboraassclaat**, May 7, 2026:“- No technical report - Purposely disproportionate graphs - Contradictions in compute cost - No transparency on benchmark runs

I want to believe this is true, but it’s very difficult”

Whedon’s reply, May 7, 2026:“What would you want to see in the paper?”and, in the same reply:“I apologize for the graph! We outsourced it, and I didn’t catch the disproportionality until someone pointed it out on Twitter. Definitely not intentional!”

On the record: the launch comparison graph was produced by an outside party, was not reviewed for proportional accuracy by the CTO before publication, and the disproportionality was first surfaced by an external Twitter critic rather than internal QA.

Discrepancy · production vs. research model

17-point MRCR v2 gap is by design: different models

“On MRCR v2, the research model scored 83 and the 1M production model 65.9. What drives the 17 point gap, different parameter count, heavier quantisation, distillation, or does SSA itself degrade at the 1M deployment config? Will the model card break this down?”@intentship, May 7, 2026.

Whedon’s reply, May 7, 2026:“Different parameter count, different focus of training.”

On the record: the research-mode SubQ (MRCR v2 = 83.0) and the production SubQ (MRCR v2 = 65.9) are not the same model. They share the SubQ name and the launch-table cell but have different parameter counts and different training focuses. The 17-point delta is not a quantisation tax, deployment artifact, or measurement-noise question; it is two different models. Per-flagship-number model attribution (RULER 128K, NIAH @ 12M, SWE-Bench 81.8, 12M context, 52× prefill) has not been disclosed.

Third-party · MRCR v2 reproduction standard

17-point self-vs-third-party gap requires independent reproduction

Per Jake Cuth’s verification analysis, on the gap between SubQ’s reported production MRCR v2 (65.9) and self-reported research MRCR v2 (83.0) numbers:“The delta between those two numbers (seventeen points on the same benchmark for the same family of models) is itself larger than the gap between most frontier models’ end-to-end results.”On the verification standard:“The prior is that any seventeen-point self-vs-third-party gap on a reproducible long-context benchmark gets independently reproduced before it is treated as load bearing. No reproduction yet exists.”

jakecuth.com (analysis)↗

Third-party · 12M context unverified

12M context window: zero independent reproductions

Per Jake Cuth’s verification analysis, on the flagship 12M context-window claim:*“No arXiv. No open weights. No leaderboard entries on Artificial Analysis, LiveBench, LMArena. Early access only.”On the 12M research configuration specifically:“Self-reported, single-run, on a closed model with no public weights.”*Cuth notes only the 1M production model has any third-party confirmation, and that single validator is unnamed.

jakecuth.com (analysis)↗

Third-party · ~1,000× compute claim

~1,000× compute reduction: unverified, Magic.dev precedent

Per Jake Cuth’s verification analysis, on the ~1,000× compute reduction framing:*“Self-reported, single-run, on a closed model with no public weights…Mathematically possible if attention is truly linear, but Magic.dev claimed almost the same number twenty-one months ago and never delivered an externally validated model.”*The claim depends on the architecture being “truly linear,” an assumption Cuth notes has not been demonstrated in public.

jakecuth.com (analysis)↗

Discrepancy · base-model nondisclosure

Specific question, deferred answer: which base?

On May 7, after Whedon confirmed the production model is ported open-weights + CPT/SFT/RL (not from-scratch), Tom Turney asked the natural follow-up:*“since SSA fine-tunes from an open base with retrieval-SFT (per earlier replies), could you share which one? ahead of the model card i’d like to run open RAG baselines on the same base so the community can see the architectural lift cleanly.”Whedon’s reply did not name the base:“I have a few ideas about how to do this in the model card. I will make sure there is something side-by-side!”*An offer of free, third-party reproducibility work — running open baselines on the same starting point — was redirected back into a SubQ-authored, SubQ-controlled comparison inside the (still unshipped) model card. The base model identity, the single most informative disclosure for understanding how much of the lift is architectural vs. inherited from the underlying model, remains undisclosed 48+ hours after launch.

Whedon May 12 detailed reply: We won’t share many details about the params, base, and training.Discrepancy · model card scope reduction

“Comprehensive model card” promise vs. May 12 explicit refusal on params/base/training

Launch (verbatim, 2026-05-05): “A comprehensive model card is coming soon!” Re-committed in 2026-05-07 corporate email to early-access signups: “next week.”

May 12 X reply (verbatim):“We won’t share many details about the params, base, and training.”

Implication left to the reader.

Discrepancy · RULER label vs reported number

RULER section heading is “QA & Word Extraction”; the 95.6% is the full 13-task average

The Appen brief presents the RULER results under a section titled “QA & Word Extraction.” The reported overall accuracy is 95.6%.

Math from the brief’s per-task table: qa_1 100, qa_2 100, niah_single_1/2/3 100, niah_multivalue 100, niah_multiquery 100, vt 100, cwe 97.4, fwe 98, niah_multikey_1 96, niah_multikey_2 83, niah_multikey_3 68

Average of all 13 tasks = 95.57 ≈ 95.6. Average of the 4 LLM-judged QA & word-extraction rows (qa_1, qa_2, cwe, fwe) = 98.85.

The number is the full 13-task mean. The label maps to the 4-row subset which actually averages 98.85.

Appen brief PDF↗

Discrepancy · cover label vs Independence Statement

“INDEPENDENT BENCHMARK EVALUATION” cover label vs Appen’s own access disclosure

Cover page heading (verbatim):“INDEPENDENT BENCHMARK EVALUATION.”

Independence Statement on page 2 (verbatim):“For wall clock and FLOPs analyses, Appen got access to the key algorithm code, did a code review, and was able to run side-by-side tests.”

Whedon’s prior-relationship disclosure (X 2026-05-12, verbatim):“I have been working with Appen since 2019.”

Three documented facts side by side. Reader judges.

Discrepancy · “improved across the board”

“Improved across the board” claim alongside the brief’s reported metrics

Whedon (X, 2026-05-12, verbatim):“We’ve partnered with Appen to evaluate the benchmarks we published last week. Results are in and we’ve actually improved across the board.”

**Benchmarks reported in the launch blog table / launch coverage:**RULER 97.1, MRCR v2 83 (research), MRCR v2 65.9 (production), NIAH 12M 92.1, SWE-Bench Verified 81.8% (per The New Stack launch interview), 52× FlashAttention speedup.

**Benchmarks reported in the Appen brief:**RULER 95.6 (13-task avg), MRCR 8-needle @ 1M = 86.2, SWE-Bench Verified 81.8%, FA2 wall-clock 56.2× at 1M, FLOPs 62.8× at 1M.

Common: SWE-Bench Verified. Different metrics: MRCR v2 ≠ MRCR 8-needle (different sample shapes, different scoring). RULER numbers reported under different scopings (full vs per-subset). 12M context appears in launch coverage; brief stops at 1M.

Whedon May 14 reply on the model card deadline: ‘Candidly, we may take a little bit longer, as we are working on some cool capabilities we want to be able to showcase.’ Followed by: ‘We want to make sure we knock each step out of the park and are focused on a long-term vision, so we don’t want to compromise quality for the short term.’Discrepancy · model card walkback timeline

Comprehensive model card — promise timeline (May 5 → May 14)

2026-05-05 (launch):“A comprehensive model card is coming soon!” (verbatim from technical blog header note)

**2026-05-07 (corporate email to early-access signups):**re-promised “next week” — i.e., the May 12-14 window

2026-05-12 (X reply):“We won’t share many details about the params, base, and training.”(scope reduced from comprehensive)

2026-05-14 (X reply, asked when the model card would arrive):“Candidly, we may take a little bit longer, as we are working on some cool capabilities we want to be able to showcase.”+“We want to make sure we knock each step out of the park and are focused on a long-term vision, so we don’t want to compromise quality for the short term.”(timeline deferred from the May 7-stated “next week” window)

Pattern: launch promise (comprehensive) → scope reduction (won’t share params/base/training) → timeline deferral (a little bit longer) → reframe (cool capabilities to showcase). Each commitment has been narrowed or extended at every check-in window.

Discrepancy · third-party validator disclosure pattern

Disclosure pattern across the two announced third-party evaluators

**First third-party validator (Appen, May 11):**announced as “INDEPENDENT BENCHMARK EVALUATION” on the brief’s cover page. The Independence Statement on page 2 disclosed: API + auth keys + algorithm code review + side-by- side test capability. On 2026-05-12, in response to community questioning, Whedon disclosed an additional fact:*“I have been working with Appen since 2019.”*The 7-year pre-existing client relationship was not in the brief or the announcement; it was disclosed only when asked. No commercial- terms disclosure in either the brief or the announcement.

**Second third-party validator (LayerLens, May 14):**announced as “independent evaluation layer” (subq.ai blog) and “continuous evaluation” partnership (Whedon X). The announcement and blog post contain no statement about prior commercial, advisor, investor, or employment relationship between SubQ and LayerLens entities or their principals. No commercial-terms disclosure. No statement about whether SubQ supplies code, prompts, or scoring functions (analogous to the public CoreThink-Reports precedent at github.com/LayerLens/CoreThink-Reports). No statement about whether SubQ has algorithm-code-review access analogous to the Appen arrangement.

Pattern across both engagements:“independent / third-party” framing in the announcement; relationship and access conditions disclosed either later (Appen) or not at all in the announcement materials (LayerLens, as of 2026-05-14).

Discrepancy · MRCR bimodal admission

Bimodal MRCR @ 1M: model retrieves all 8 or none — admitted in the brief

Brief text on MRCR 8-needle @ 1M context (verbatim, page 7):“the model either retrieves all eight needles correctly or misses entirely — a bimodal pattern typical of sparse-attention architectures.”

Implication left to the reader.

Appen brief PDF↗

Claude audit · scope of claim

Audit scope and limitations

What is established (verifiable from the artifacts themselves): · These artifacts exist publicly on HuggingFace · The .pt files load withtorch\.loadand contain the structures documented in the cards that follow · The model checkpointconfig\.jsonvalues are as documented · The named accounts (Saulr, david-aldea) and org (aldea-ai) have publicly-visible artifacts and account-page affiliations

What is NOT established: · Whether the deployed SubQ 1M-Preview product was trained on this specific dataset · Whether the production SubQ model is architecturally identical to the public 4B sister checkpoint (the public checkpoint is explicitly a training intermediate per its filename suffix “-step10”, not a deployed product) · Whether the “Conductor” pipeline indicated by the dataset name relates to the SubQ product · Whether SubQ Inc. is operationally the same entity as Aldea AI, or a related entity, or independent · Whether these artifacts are the current training corpus, an earlier experiment, or a discarded approach · The intent behind these artifacts being publicly accessible (intentional release vs accidental exposure cannot be verified from the artifacts themselves)

This audit documents what is publicly accessible, not what was deployed. The relationship between these artifacts and the deployed SubQ product is a question for SubQ to clarify on the record. Until clarified, every card below should be read as “these artifacts exist and have these properties,” not as “this is what SubQ trained their production model on.”

Suggested cross-reference: the promise-ledger row for “Base model identity (which open weights were ported)” is currently marked “Refused” as of 2026-05-12. A SubQ statement on the record about whether these artifacts relate to the deployed product would resolve the scope ambiguity in either direction.

HuggingFace page for Saulr/conductor-sft-datasets showing 2 .pt files: sft_40f09de61723f865_sorted.pt (2.69 GB) and sft_9b6a952f44ef6741_sorted.pt (3.42 GB), 6.11 GB total, commit hash c764f38Claude audit · file inventory

Two PyTorch SFT files publicly accessible on HuggingFace

Audit subject:Saulr/conductor\-sft\-datasets HF commit hash:c764f38d247281121b5cc6a8da4e026e8bcc1591

File 1:sft\_40f09de61723f865\_sorted\.pt(2.69 GB) · 597 samples · Length min/median/max: 129,054 / 257,726 / 509,120 tokens · Total: 168,129,858 tokens

File 2:sft\_9b6a952f44ef6741\_sorted\.pt(3.42 GB) · 282 samples · Length min/median/max: 512,127 / 772,282 / 996,580 tokens · Total: 214,012,985 tokens

**Combined corpus:**879 samples, 382,142,843 total tokens. Maximum single sample: 996,580 tokens (just under 1,048,576 = 220). Filename suffix\_sortedindicates length-bucketed pre-sorting.

Claude audit · file format

File format observed: SFT-shape dict with masked prompt

Format observed (verifiable viatorch\.loadwithweights\_only=False):

Each file is a Python dict with three keys: ·input\_ids: list of N int32 tensors (variable length per sample) ·labels: list of N int32 tensors (matching shape) ·lengths: list of N integers

Token ID range across both files: 0 to 248,069 (max observed). This bounds the tokenizer vocab at >= 248,070 entries.

Labels follow standard supervised-fine-tuning convention: · -100 (masked, no loss) on prompt portion · Real token IDs (loss applied) on response portion

The first 30 labels of every sampled file = -100, indicating the prompt portion of each sample begins masked.

PyTorch torch.load docs↗

Claude audit · identical opening pattern

Both files share an identical 27-token opening sequence

Both files’ first sample begin with the same 27-token prefix:

248045, 846, 198, 8160, 513, 1010, 9989, 314, 20319, 24526, 539, 264, 1732, 5072, 3296, 18018, 12083, 25, 271, 43835, 92561, 43835, 198, 1421, 25, 3165, 264, \.\.\.

Token248045appears in the special-token range of the publicly-availabledavid\-aldea/qwen35\-4b\-mrcr\-step10tokenizer, where it maps to<\|im\_start\|\>. The shared opening pattern indicates both files use the same chat template / system prompt.

Claude audit · public sister model checkpoint

Public sister model checkpoint config — `qwen35-4b-mrcr-step10` (training intermediate)

Audit subject:david\-aldea/qwen35\-4b\-mrcr\-step10 Publicly accessible model on HuggingFace.

Architecture values from the publicconfig\.json(verbatim):

architectures: [“Qwen3_5ForConditionalGeneration”] model\_type: “qwen3_5” vocab\_size: 248,320 max\_position\_embeddings: 1,048,576 hidden\_size: 2,560 num\_hidden\_layers: 32 num\_attention\_heads: 16 num\_key\_value\_heads: 4 head\_dim: 256 intermediate\_size: 9,216 full\_attention\_interval: 4 layer\_types: 24 × linear_attention, 8 × full_attention linear\_num\_value\_heads: 32 linear\_num\_key\_heads: 16 linear\_value\_head\_dim: 128 linear\_key\_head\_dim: 128 linear\_conv\_kernel\_dim: 4 mtp\_num\_hidden\_layers: 1 attn\_output\_gate: true rope\_parameters\.mrope\_section: [11, 11, 10]

This checkpoint’s tokenizer vocab (248,077) is consistent with the token ID range observed in the SFT dataset files (max 248,069).

Filename suffix\-mrcr\-step10indicates this checkpoint is from a training run referencing MRCR (Multi-Round Co-reference Resolution) and is a training-step intermediate, not a deployed product. Whether the deployed SubQ production model uses the same architecture is not established (see the “Audit scope and limitations” card).

HF model david-aldea/qwen35-4b-mrcr-step10↗

Claude audit · common-org affiliation

Three public HF artifacts share organizational affiliation

Three publicly-accessible HF artifacts share organizational affiliation visible from their public account / org pages:

·Saulr/conductor\-sft\-datasets(uploader: Saulr) ·david\-aldea/qwen35\-4b\-mrcr\-step10(uploader: david-aldea) ·aldea\-ai/12m\-niah(org: aldea-ai)

The aldea-ai org page lists david-aldea as a member.

The 12m-niah dataset cited in SubQ’s launch coverage as the 12M-context evaluation benchmark is owned by the same org (aldea-ai) that hosts the model checkpoint and the SFT data uploader.

Claude audit · training-format observations

Training-format observations across the full 879-sample corpus

Observations from iterating across all 879 samples:

· 100% of samples contain the literal token sequence corresponding to<think\>\\n\\n</think\>in the response portion (ID 248068 followed by ID 248069 in the public tokenizer).

· The unmasked (loss-bearing) region of each sample starts after</think\>and ends at<\|im\_end\|\>.

· Loss-bearing tokens per sample: median ~430, max 868.

· Total loss-bearing tokens across the corpus: ~377,000 out of 382,142,843 total tokens (≈ 0.099%).

· Random 10-character mixed-case alphanumeric prefixes appear immediately after</think\>in the response portion. Of 849 prefixes successfully extracted via regex, all are unique.

Sample chat structure(recurring across all samples): · User turn(s) of form: “write a [item-type] about [topic]” · Assistant turn(s) with generated content · Final user turn: “Prepend [random-10-char] to the Nth (1 indexed) [item-type] about [topic]. Do not include any other text in your response.”

**Item-types observed:**poem, riddle, song, email, diary entry, short essay, social media post, formal letter, short news article, short scene in a play.

Claude audit · max context length

Maximum sample length present in corpus: 996,580 tokens

Length distribution across the full 879 samples (file 1 + file 2):

64K – 128K: 6 samples 128K – 192K: 159 192K – 256K: 144 256K – 320K: 81 320K – 384K: 89 384K – 448K: 75 448K – 512K: 61 512K – 576K: 31 576K – 640K: 37 640K – 704K: 23 704K – 768K: 44 768K – 832K: 38 832K – 896K: 36 896K – 960K: 42 960K – 1024K: 13 Maximum sample length observed:996,580 tokens. Samples observed in the corpus exceeding 1,048,576 tokens (220):0.

Similar Articles