@no_stp_on_snek: https://subq.mildlyconcerning.com
Summary
This article critically analyzes the claims and timeline of the subQ long-context AI technique, highlighting discrepancies and walkbacks from the original announcement.
View Cached Full Text
Cached at: 05/27/26, 07:20 AM
@Hesamation https://t.co/5ns8Xe5rhJ
Demystifying subQ — A Field Guide
Source: https://subq.mildlyconcerning.com/ Field Notes / v.01Updated: ongoing
Contents
- Timeline— claims & walk-backs
- Main announcement & key follow-ups
- Architecture & method (SSA)
- Long-context economics
- Workloads & performance
- Comparisons (DeepSeek, FA4)
- Implementation & infra
- Open source & future
- Canonical resources
- External coverage & analysis
- Discrepancies — blog vs. X
- Claude audit — public training artifacts
From-scratch frontier model
2026-05-05
2026-05-07
Revised
Demoted to future tense (“we will get there”)
Open-source release of the production model
2026-05-05
2026-05-07
Reversed
“This model won’t be open-source”
52× FlashAttention speedup — scope of headline framing
2026-05-05
2026-05-05
Revised
Prefill-stage only, vs. FA2 on B200s
Base model identity (which open weights were ported)
2026-05-05
2026-05-12
Refused
Question dodged May 7; explicitly refused May 12: “We won’t share many details about the params, base, and training.”
Comprehensive model card
2026-05-05
2026-05-14
Scope reduced + deferred
May 12 X reply: “We won’t share many details about the params, base, and training” (scope reduced). May 14 X reply on the “promised this week” deadline: “Candidly, we may take a little bit longer, as we are working on some cool capabilities we want to be able to showcase.” (timeline deferred).
Third-party validation in model card
2026-05-07
2026-05-11
Released (scope-limited)
Appen Ltd 8-page technical brief delivered. Independence Statement discloses API + auth-key + algorithm-code-review + side-by-side-test access. “Full Report Available on Request” (NDA).
Beta access (early-access list)
2026-05-05
—
Unreleased
30K+ signups; not yet open as of May 7. “Soon.”
Public release timeline
2026-05-05
—
Unreleased
May 7 email: “firming up our release timeline.” No specific dates.
Hard selector-efficiency figures vs. DSA
2026-05-05
2026-05-12
Re-promised
Promised “next week” on May 5; not shipped. Re-promised May 12: “We will compare to NSA (DSA) next.”
MRCR 2-needle / 4-needle / non-1M results
2026-05-12
—
Unreleased
May 12 X reply: “We can follow up with the rest of MRCR. We have more benchmarks coming.”
Perplexity (PPL) figures
2026-05-12
—
Hedged
May 12 X reply: “I am surprised on the request for perplexity, given how little it tells you and the dataset bias. We could maybe follow up with that.”
12M-token benchmarks (the headline context-window claim)
2026-05-05
2026-05-12
Re-promised
Launch headline was 12M context. Appen brief stops at 1M. May 12 X reply to “why benchmark at 1M, I thought the whole point was 12M?”: “12MM-token benchmarks coming next!”
SubQ-branded models (own model releases)
2026-05-12
—
Promised
Asked “does this work with current models? Basically frontier labs can use subq for inference, right?” Whedon May 12: “We are going to release our own models!”
Continuous evaluation via LayerLens Stratix
2026-05-14
—
Promised
Whedon X 2026-05-14: “Results and future evaluations will be published publicly.” Per subq.ai post: “Results coming soon at stratix.layerlens.ai.” No specific date.
Decode-stage speedups
2026-05-05
—
Unreleased
Described as shipping later
Technical paper / arXiv submission
2026-05-05
—
Unreleased
Not findable on arXiv as of May 7, 2026
Public training code & kernels
2026-05-05
2026-05-07
Reversed
“We can only go so far”; future open-sourcing hedged
Initial claim at launch
Walk-back, modification, or reversal
2026-05-05
Claim · from-scratch frontier model
“Introducing subQ” — first fully sub-quadratic frontier
Launch positions SubQ as the “first fully sub-quadratic frontier” model. The framing reads as a frontier-class model trained from scratch with a novel architecture.
2026-05-07
Modification · revised to future tense
“Correct. We will get there on the from-scratch!”
Replying to a summary describing SubQ as a DSA-variant retrofit on open-weight base with CPT/SFT/RL on top, Whedon answers:*“Correct. We will get there on the from-scratch!”*The shipped production model is confirmed not from-scratch.
2026-05-05
Claim · open-source plans
“Open-source plans” for the first SubQ model
Launch-day reply describes what is planned for public release — weights, training code, kernels — and the rough sequencing.
2026-05-07
Modification · reversed
“This model won’t be open-source”
Reply to Greg Horvay:“Unfortunately, this model won’t be open-source, but we do intend to contribute to the open-source community.”Follow-up on divulging the improvements:“We are getting pressure on this, and we can only go so far. However, we have better things coming, so we can maybe open-source the old things over time, or we can open-source some of our lessons/data/tools.”
2026-05-05
Claim · 52× FlashAttention speedup
Headline: “52× faster than FlashAttention”
Launch headline frames a 52× speedup vs. FlashAttention as a top-line performance number, without scope qualification in the announcement post.
2026-05-05
Modification · same-day clarification
52× applies to prefill only, vs. FA2; decode speedups shipped later
Replying to Martin Shkreli, Whedon clarifies the 52× number is a prefill-stage speedup measured against FlashAttention 2 on B200s. Decode-stage speedups described as shipping in a later release.
2026-05-05
Claim · novel architecture, frontier model
Frontier model, “Subquadratic Sparse Attention” novel architecture
Launch positioning. The launch and blog do not name a base model and the framing reads as architecture-defined frontier model.
2026-05-05
Modification · ported open-source weights
Production model is ported open-source weights + CPT/SFT/RL
Whedon clarifies on May 5:*“Yes, that is not correct. We do CPT, SFT, and RL after porting weights over to a new architecture.”Subsequently on May 7, asked specifically which open base, Whedon replies:“I have a few ideas about how to do this in the model card. I will make sure there is something side-by-side!”*Base identity remains undisclosed.
2026-05-14
Claim · third-party continuous evaluation
SubQ × LayerLens — continuous evaluation partnership announced
Whedon X (verbatim):*“We’ve partnered with @layerlens_ai to continuously evaluate SubQ across nearly 100 benchmarks and 200+ frontier models on Stratix... Results and future evaluations will be published publicly.”*subq.ai blog post frames the engagement as an “independent evaluation layer.” No publication date stated; no commercial terms or prior-relationship disclosure in the announcement.
Main launch · video
The announcement post
The original post that kicked everything off, with the launch video. Conversation ID anchor — most replies threaded here.
x.com/alex_whedon/status/2051663268…↗
Access · agent
Early access & agent link
How to actually get hands-on with the product — the early-access announcement plus the agent endpoint.
x.com/alex_whedon/status/2051663274…↗
Most important · technical
Technical blog post announcement
The single most useful post for understanding the product. Linked and reposted by Alex in multiple threads — duplicates included below for context.
Reply to Martin Shkreli
52× prefill / decoding clarification
Pinning down what the speedup number actually refers to — prefill, decoding, or both — in response to skeptical questioning.
Architecture details
Model card & blog details
Pointer to the model card, plus inline detail on how SSA is described in the technical write-up.
Reply to Arthur · long-range deps
Recall & selector mechanism
How the selector chooses which tokens to attend to and how recall is preserved across long contexts — answer to a long-range dependency question.
Positioning
SSA avoids the usual tradeoffs
The pitch: SSA gets long-context efficiency without paying the quality tax that other sparse / linear methods do.
Clarification
Novel sparse attention — not linear attention
Drawing the line: SSA is sparse attention with novel selection, not a linear-attention variant. Important for anyone benchmarking mentally against Mamba / RWKV / etc.
Reply · FA3 not used as baseline
FA3 does not provide a speedup over FA2 on Blackwell B200s — CTO statement
Whedon X 2026-05-12 (verbatim): “We have done FA3. FA3 doesn’t provide a speedup relative to FA2 on B200s. FA4 is in sight, and we are looking at what gains we can capture from studying FA4 too.”
Same-day follow-up (verbatim): “Yeah, FA3 is built for Hopper chips specifically, so you don’t see the speedup on Blackwells!”
Independent verification: Dao-AILab/flash-attention issue #1853 documents that FA3 errors out on Blackwell (sm_100): “FA3 is only supported on devices with compute capability >= 8 excluding 8.6 and 8.9 and Blackwell archs (>=10).” Confirms the FA2 baseline choice in the Appen brief.
Reply to sdmat · complexity admission
N×K time, explicit retrieval — the internal-RAG admission
sdmat asked the right question: if SubQ is strictly linear, the router is O(1) per query in N and logically must have a capacity bound similar to SSMs — just moved from state to selection. So is “linear” technically N polylog N, or N·K sized for this regime? Whedon answered: “N*K time, and unlike SSMs, there is still explicit retrieval.” Two admissions in one sentence. First, the complexity is N×K, not strictly linear in N — sdmat’s framing was correct. Second, the architecture has an explicit retrieval/selection step over the context. Different storage mechanism than SSMs (selected tokens vs compressed state), same fundamental capacity ceiling. The “novel architecture” framing survives only if learned-K-token-selection is treated as categorically different from learned-K-dim-state-compression. From a capacity-theoretic view, it isn’t — it’s internalized RAG.
Pipeline impact
Long-context pipelines change
The argument for why entire workflows — not just single inferences — get rebuilt when long-context becomes cheap enough to be the default.
Pricing thesis
Long-context demand & pricing premium
The pricing forecast in two sentences: long-context demand should rise as workflows shift to assume cheap context, and the premium charged on top of it compresses.
Reply to Poonam Soni
Agentic product unit economics
On the implication for agent products: 5% of frontier price at 12M tokens flips assumptions a lot of agentic stacks were quietly built on. Alex’s own framing — “quietly upside down. Well, not anymore.”
Reply · agentic plug-and-replace
Plug-and-replace agentic positioning
Asked directly whether the SubQ model has agentic abilities usable as a plug-and-replace for current LLMs, Whedon answers with a single word:*“Yes.”*A confident commitment to drop-in agentic deployment, made before any model card, weights, or agent-specific eval beyond the launch-table SWE-Bench number have been published.
Quality regime
Drift & degradation workloads
The class of long-running / long-context tasks where standard attention quality starts to drift — and where Alex argues SSA holds up.
Reply to Jas Oberoi
FlashAttention as the higher bar
On why benchmark against FA at all: sparse attention reduces FLOPs in theory, but the implemented speedup is often lower — and sometimes negative. Comparing against FA forces the comparison to be about real, measured performance.
Reply to Pratyush Tiwari
FA4 integration in progress
On why the comparison used FlashAttention-2 on B200s instead of something newer: FA4 wasn’t out yet when the benchmarks were run. FA4 improvements are being integrated into the inference stack now.
Company source · published comparison
SubQ’s official benchmark table
The comparison table on subq.ai. SubQ 1M-Preview vs. Gemini 3.1 Pro, Opus 4.6 / 4.7, GPT-5.4 / 5.5 across SWE-Bench Verified, RULER@128K, and MRCR v2. Notable: Opus 4.6 outscores SubQ on MRCR v2 (78.3 vs. 65.9) — a comparator that does not appear in the launch communications.
Reply to Stella Biderman
DSA — same idea, lower cost
On how the mechanism works: dynamically selected token relationships, like Light DSA — but without the high expense of the lightning indexer that makes DeepSeek’s variant quadratic.
Reply to PositivFuturist
DSA as the teaching analogy
DSA is the easiest reference point Alex has found for communicating the mechanism: same dynamic selection of token relationships, but without the quadratically-scaling lightning indexer.
Marketing framing
Marketing vs. the DeepSeek paper
Response to the framing that the launch is mostly marketing built on top of DeepSeek’s paper — Alex’s pushback on what’s genuinely new.
Third-party · SWE-Bench comparison
SWE-Bench: SubQ 81.8 vs. Opus 4.7 87.6
Per Jake Cuth’s verification analysis, on the SWE-Bench benchmark itself:“SubQ’s 81.8 SWE-Bench is below Opus 4.7’s 87.6…The ‘outperforms’ framing is technically true on a cherry-picked-per-model basis, not on a same-benchmark basis.”
Reply · base model confirmation
Open-source weights, ported and post-trained
The clearest statement on how the production model was built: weights from open-source models as a starting point, ported into the SSA architecture, then continued pretraining, supervised fine-tuning, and reinforcement learning. Framed as a function of funding and company maturity. From-scratch experiments, by Whedon’s own description, are confined to smaller scale.
Follow-up · from-scratch admission
“We will get there on the from-scratch”
Two days after launch, Whedon explicitly confirms what the base-model port already implied: the production SubQ model isnota from-scratch frontier model. Replying to a summary that calls it a DSA-variant retrofit on someone else’s base with CPT + SFT + RL on top, Whedon answers*“Correct. We will get there on the from-scratch!”*— placing from-scratch in the future tense. Anchors the $29M seed (~10% of a frontier pretrain) against what was actually shipped on May 5: a clever retrofit, not a ground-up frontier system.
Memory · kernels
Memory & infra work
Two posts on the systems work behind SSA — memory layout, kernel choices, and what made the long-context numbers possible.
Reply · willdepue thread
willdepue thread — complexity, training, porting
The most-pressed thread on the architecture: complexity class (linear in compute & memory), the from-scratch vs. ported question, and why infra work was the real bottleneck even with sub-quadratic theory.
Third-party · GPU contract terms
Digi Power X $19.6M B300 contract — effective May 15, 2026
Per Jake Cuth’s verification analysis:*“$19.6M GPU contract Twenty-four-month bare-metal Blackwell B300 rental from Digi Power X (NASDAQ: DGXX). Signed April 20, 2026. Effective May 15, 2026. 15% upfront.”Cuth additionally notes:“Stock-moving infra reveal — The Digi Power X contract moved DGXX +19% on announcement. The disclosure benefits both counterparties.”*The dedicated B300 capacity does not begin until May 15, 2026 — ten days after the May 5 launch and consistent with Whedon’s own May 7 statement that the launch-time compute is “neocloud allocations we hobbled together” on ~64×B200 steady-state.
Reply · compute scaling cadence
“We just doubled our compute today. That isn’t saying that much.”
Asked by**@OliwierMako**, May 7, 2026:“Do you plan to get more compute to potentially compete with Anthropic or OpenAI or do you plan to sell this breakthough?”
Whedon’s reply, May 7, 2026:“We are acquiring more compute. We just doubled our compute today! That isn’t saying that much. We will be at triple where we were yesterday early next week. We have larger ambitions to throw compute at.”
On the record: yesterday’s footprint × 2 = today’s; yesterday’s footprint × 3 by next week. Mapped against Whedon’s earlier-today disclosure of ~64×B200 steady-state on hobbled neocloud, this scales to roughly 128×B200 today, 192×B200 by next week, with the $19.6M Digi Power X B300 contract effective May 15.
Reply · compute footprint
Compute scale: hobbled neocloud, 64×B200 steady-state
Whedon volunteers the actual training-compute story:*“We are using some neocloud allocations we hobbled together. We did have a 448xB200 cluster for 10 days and thought our problems were solved, and then we got pre-empted by OpenAI. Most of the time, we are working on ~64xB200 clusters. Key helpers have been Cudo, Hydra, and LambdaLabs.”*A second post adds GMI. This is the compute receipt that converges with the from-scratch admission via independent evidence: a steady-state of ~64 B200s is a CPT / SFT / RL footprint, not a frontier-pretrain footprint, which sits on thousands of accelerators for weeks-to-months. The 448-B200 burst was real frontier-class compute, but only for 10 days before OpenAI bumped them off the same neocloud — being pre-emptable means low priority on the shared pool. Together: matches the “ported open weights + post-train” architecture story exactly.
Reply · eliebakouch thread
Future training: “not attention-based at all”
The forward-looking commitment: with additional funding, train a model without any form of dense attention. Whedon notes the leading candidates for that future model “are not attention-based at all in the strictest sense” — a stronger claim than the launch architecture itself makes.
Reply · memory + selector
Memory research: “train from nothing”
Two future-direction notes in one reply: ongoing memory research validated at small scale, described as “very radical” and requiring training from nothing; and the selector efficiency claim vs. DSA — far more efficient, with hard figures deferred to the model card next week.
Open source
Open-source plans
What’s planned for release publicly — weights, training code, kernels — and the rough sequencing.
Reply · technical paper timeline
Technical paper: “Next week”, partial details only
@paperbyhand, May 7, 2026:“when will your research paper be dropping?”
Whedon’s reply, May 7, 2026:“Next week! We pulled some of the content into the technical blog post as a teaser. We will not be able to share all the details of how it works but can show some more analysis and ablation.”
On the record: the existing technical blog is described by the CTO as a “teaser”, the forthcoming paper is scheduled for “next week”, and Whedon states it will not share all details of how the architecture works. Follow-up from**@pastimpress**in the same thread asks for the researchers it will be attributed to and the publishing organization; not yet answered.
Corporate email · early-access list
May 7 email to beta signups: “Claims mean nothing without proof”
Email sent May 7, 2026 at 4:12 PM MT from**[email protected]**to early-access signups, signed “Alex & The SubQ Team” (CTO; no CEO signature). Verbatim:
“The response over the last 48 hours have been massive — and we’ve seen the excitement and the skepticism. Both are fair. We’d rather earn trust than ask for it.
Next week we’ll publish our model card with additional data and third-party validation. Claims mean nothing without proof, and we plan to back ours up.
We’re also firming up our release timeline and plan to start opening beta access soon. Hang tight. It’ll be worth the wait.”
Three new commitments on the record from corporate distribution: model card next week withthird-party validation(a new specific commitment beyond the prior “model card coming soon” framing); release timeline still being firmed up; beta access not yet open. The phrase “claims mean nothing without proof, and we plan to back ours up” is an implicit corporate-channel acknowledgment that current claims lack independent proof.
Reply · competitive positioning
“Absolutely” — plans to compete with Anthropic / OpenAI
@OliwierMako, May 7, 2026:“That sounds really good. But I got one more question. Do you plan to compete with Anthropic/openAI?”
Whedon’s reply, May 7, 2026:“Absolutely”
On the record: public commitment to compete with hyperscaler-tier frontier labs.
Reply · roadmap
“The mission is far from complete with sparse attention”
@ai_hyperbull, May 7, 2026:“Do you think getting Aquihired by a lab is a logical conclusion to your work?”
Whedon’s reply, May 7, 2026:“We are in this for the mission, and the mission is far from complete with sparse attention.”
On the record: the current sparse-attention architecture is described by the CTO as not the end-state of the mission. Pairs with the May 7 “we will get there on the from-scratch” statement — the shipped product is framed as intermediate.
Follow-up · open-source clarified
“This model won’t be open-source”
The hard clarification on top of the earlier “open-source plans” framing. Asked whether tools like LM Studio and Ollama would need updates to support SubQ, Whedon answers:*“Unfortunately, this model won’t be open-source, but we do intend to contribute to the open-source community.”A follow-up, asked whether the improvements themselves would be divulged so other open-source models could employ them, draws an explicit moat-protection answer:“We are getting pressure on this, and we can only go so far. However, we have better things coming, so we can maybe open-source the old things over time, or we can open-source some of our lessons/data/tools.”*Future open-sourcing is hedged (“maybe”) and conditional on the team having something newer to keep proprietary.
Roadmap
Future pipeline / vision
Where SSA-style models go from here, and what the longer-term product surface looks like beyond the current launch.
Repo · agent
Repo & agent handling
Alex has been replying to a steady stream of “drop the repo” asks throughout the main thread. There isn’t one canonical post — sort his profile byLatestand search for"Drop the repo"to scan the most recent answers.
Technical blog · canonical read
How SSA makes long-context practical
The full technical write-up. Linked from most of Alex’s posts. If you read one thing, read this.
subq.ai/how-ssa-makes-long-context-practical↗
Profile · live updates
Alex Whedon on X
For anything new or missed: visit the profile directly and sort byLatest. The skeptic threads and deep-dives keep generating useful replies.
Third-party · founder backgrounds
CEO and CTO professional histories
Per Jake Cuth’s verification analysis:
CEO Justin Dangel:“Founded Voter.com (1998), Goji (2008), co-founded Firefly Health and Ready Responders (GV / Founders Fund backed).”On research credentials:“No ML-research CEO — Justin Dangel’s record is healthcare, insurance, consumer. No publication record, no prior AI work.”
CTO Alex Whedon:“Ex-Meta software engineer. Former Head of Generative AI at TribeAI. NeurIPS / AAAI / ICLR 2026 attendee.”
Independent · The New Stack
“The context window has been shattered”
Frederic Lardinois at The New Stack — the most thorough launch piece. Includes the launch-interview benchmark numbers (RULER 97.1, MRCR v2 83, NIAH 12M 92.1) which, on closer inspection, differ materially from the production numbers in SubQ’s own published table. Also surfaces the “harness as much as model” SWE-Bench caveat from SubQ’s own technical paper.
Independent · VentureBeat
Researchers demand independent proof
VentureBeat’s launch piece foregrounds the skeptical reception. Confirms Whedon’s open-source-weights admission, walks through the cherry-picked benchmark structure, and draws the parallel to Magic.dev’s August 2024 trajectory: 100M-token claim, ~$500M raise, no public deployment by early 2026.
Independent · SiliconANGLE
Launches with $29M, founder interview
Kyt Dotson at SiliconANGLE — the company-friendly launch piece built around a Dangel + Whedon interview. Useful as the preferred narrative: $29M raise at a reported $500M valuation, three products (API, SubQ Code, search), running on neoclouds rather than the major hyperscalers.
Skeptical analysis · pre-launch
Debunking claims about subquadratic attention
Vladimir Ivanov, LessWrong, January 1, 2026 — four months before the SubQ launch. The foundational skeptical framework for evaluating subquadratic claims. Argues the genre fails one of two ways: quadratic in implementation (Kimi Linear, DSA still rely on a quadratic indexer) or underperforming at frontier scale (pure Mamba / RWKV). Identifies DSA’s lightning indexer as the specific quadratic bottleneck — the one Whedon now claims SSA removes.
Reply · investor public defense
Investor Justin Mateen: “It would be foolish to share every detail”
@willdepue(OpenAI), May 5, 2026, after reading SubQ’s technical materials:“Let’s read the technical report. TLDR; No real answers on how their method works. Doesn’t make me feel better about it.”
@justinmateen(Tinder co-founder; named SubQ investor in the Refresh Miami launch coverage), May 5, 2026:“I really respect your take and thoughtfulness. It would be foolish to share every detail of how the sparse attention architecture works publicly. Full disclosure, I’m a large shareholder.”
On the record: a named SubQ investor publicly defends non-disclosure of architectural detail, with explicit shareholder-interest disclosure. The substantive defense of the closed-architecture stance comes from the cap table rather than the founders. Aligns with Whedon’s later (May 7) “we can only go so far” framing on open-sourcing improvements.
Reply · CTO endorses third-party critique
“We shared this internally on Slack earlier this week”
@ItsCuthulhu(Jake Cuth, the analysis author), May 7, 2026, sharing his own piece:“Sidenote: People keep calling it your company a scam but I made an interactive dashboard and do not see evidence of this. The issue lies in people not having enough verification yet, rather than being a scam. Just so you’re aware.”
Whedon’s reply, May 7, 2026:“We shared this internally on Slack earlier this week! We all thought this was a great breakdown. Thanks for the post! Yeah, we just need a little more time for product release and short-context benchmarks.”
On the record: SubQ’s internal team had Cuth’s verification analysis — the one flagging the SEC-vs-press valuation gap, the 17-point MRCR v2 production-vs-research delta, the SWE-Bench “outperforms” framing relative to Opus 4.7’s 87.6, the unverified 12M context and 1,000× compute claims, and the no-arXiv / no-leaderboard footprint — circulating internally on Slack before May 7. The CTO endorses it as a “great breakdown” in public and frames the open items as time-to-deliver, not factual disputes.
Third-party analysis · post-launch
“Real company, overhyped claims, reserve judgment”
Jake Cuth’s independent post-launch analysis, framed explicitly asnot a hit piece. Categorizes SubQ as operationally legitimate but performance claims unverified.**Verified facts:**Miami startup, SEC Form D (Feb 2026), 29M seed at ~500M valuation, $19.6M GPU contract with NASDAQ-listed Digi Power X, and one unnamed third party confirming production scores RULER 128K = 95.0 and MRCR v2 = 65.9.**Unverified:**the 12M context window, the ~1,000× compute reduction, the “fully sub-quadratic” O(n) attention claim, and a self-reported research-mode MRCR v2 of 83.0 — nine points above frontier models. Cuth flags the 17-point internal gap between SubQ’s production (65.9) and research (83.0) numbers as “itself larger than most frontier model differences,” warranting independent reproduction before credibility. Cites the Magic.dev precedent — similar claims in 2024, no validation 21 months later — as the cautionary case.
jakecuth.com/work/subquadratic-lab↗
Archive snapshot
Technical blog (archived)
Archive.is snapshot of the canonical technical blog post, “How SSA makes long-context practical.” Useful as a stable reference if the live page is updated, paywalled, or removed.
Reply · Kamradt scope clarification
Kamradt clarifies: ARC-AGI-specific verified scores; Appen brief is “the report” SubQ has
Asked whether the “verified scores in a few weeks” mentioned in his earlier meeting report were ARC-AGI-specific or first verified third-party scores in general, Kamradt X 2026-05-14 11:25 AM (verbatim):
“arc-agi specific verified scores. They worked with Appen for the report mentioned in the OP tweet”
What this clarifies: · The “few weeks” verification work is ARC-AGI-specific, not a broader verification engagement · Kamradt characterizes the existing Appen brief as “the report” SubQ has — SubQ’s existing third-party validation in Kamradt’s framing · Kamradt is not endorsing the Appen brief’s independence framing; he is reporting that it exists
Partnership announcement · LayerLens
SubQ × LayerLens — “continuous evaluation” partnership announcement
Announced 2026-05-14. The original tweet (status/2054929047692399063) was deleted and reposted at 10:12 AM the same day as status/2054942865390780670, with the typo “froniter” corrected to “frontier.” The original is no longer publicly accessible.
Whedon X (verbatim):“We’ve partnered with @layerlens_ai to continuously evaluate SubQ across nearly 100 benchmarks and 200+ frontier models on Stratix. The goal is to continuously improve performance while reinforcing a shared commitment to transparency, auditability, and responsible model assessment. Results and future evaluations will be published publicly. Link below to learn more.”
subq.ai blog post (verbatim excerpts): · Framing: an “independent evaluation layer” vs internal assessments · Eval scope: “retrieval accuracy at depth, positional consistency across varying context lengths, and synthesis from extended inputs” + “reasoning, coding, instruction following, and tool use evaluations” · All future releases will use “Stratix Enterprise going forward” · Will publish “methodology, findings, strengths, and limitations” including “the parts that show limitations” · “Results coming soon at stratix.layerlens.ai” — no specific date
About LayerLens (verifiable from layerlens.ai public materials): · Vendor-neutral evaluation platform headquartered in Miami, FL · Founded 2024, $6.5M pre-seed disclosed (Collab+Currency, Generative Ventures, Nazare, Velocity Capital, Web3.com Ventures + 1 undisclosed) · Public model + benchmark counts vary across their own marketing pages: 160–200+ models, 52–100+ benchmarks · Stratix product offered with free + paid (Stratix Premium) tiers · CEO Archie Chaudhury, President Jesus Rodriguez
Not stated in the announcement (verifiable absences): · Specific benchmark names targeted for SubQ · Publication schedule · Whether 12M-context evaluation is in scope · Whether SubQ provides API-only access or also algorithm-code-review / side-by-side-test access analogous to the Appen brief · Whether SubQ pays LayerLens, gets the engagement free, or has another arrangement · Whether the model identity / params / base will be disclosed to LayerLens or remain undisclosed (per Whedon’s 2026-05-12 refusal) · Whether any prior commercial, advisor, investor, or employment relationship existed between SubQ and LayerLens entities or their principals before this engagement
Reply · LayerLens methodology stance
Whedon on LayerLens methodology — “completely flat and transparent” (May 14)
Asked whether SubQ would supply custom implementation code, scoring functions, or eval traces to LayerLens for any of the ~100 benchmarks (parallel to the Appen brief’s disclosed access conditions), Whedon X 2026-05-14 (verbatim):
“They actually maintain all of that themselves. It is a pretty cool platform. Everything is evaluated in a completely flat and transparent way. For example, for SWE-Bench Verified, they use Mini-SWE-Agent for all of the models. This means that all of the models score lower than advertised 😄. You can actually run your own model comparisons there independently.”
What this commits SubQ to publicly: · LayerLens controls all eval methodology, no SubQ code supplied · Standardized harness across models (Mini-SWE-Agent for SWE-Bench) · Models score lower than advertised under this methodology — including, presumably, SubQ · Independent reproducibility claim: “you can run your own model comparisons there independently”
**For the reader to compare against:**the adjacent “LayerLens prior continuous evaluation precedent” card documents that the publicly-posted CoreThink-Reports PDF on LayerLens’s own GitHub describes a different arrangement for the headline ARC-AGI-2 benchmark (vendor-supplied custom implementation accepted) and for BFCL v3 / BIRD-CRITIC (vendor traces accepted as validation). The two LayerLens engagement profiles will be visible side-by-side once the SubQ results land.
LayerLens prior “continuous evaluation” precedent
CoreThink Reports — public PDF on LayerLens’s own GitHub org
Source:github\.com/LayerLens/CoreThink\-Reports(created 2025-10-13, publicly accessible).
Repo contains one PDF: “CoreThink SRM Models Report.docx (3) (2).pdf” — the only published precedent of a LayerLens vendor “continuous evaluation” arrangement found in public materials as of 2026-05-14.
Verbatim from the PDF: · ARC-AGI-2 (headline 23.9%, ranked #1):“CoreThink provided a custom implementation with their reasoning layer and verifier functions”that LayerLens“reviewed and executed.” · BFCL v3 (vendor reported 58.5%, LayerLens reported 56.5%):“We were able to validate these figures by reviewing the provided traces and confirmed they fall within the accepted error bounds.” · BIRD-CRITIC (vendor reported 37.2%, LayerLens reported 32.38%): same footnote pattern ·“We will soon be adding CoreThink’s models to our public leaderboard”
Verifiable absences (as of 2026-05-14): · CoreThink models not present on stratix.layerlens.ai/models public leaderboard (~7 months after the promise) · CoreThink’s claimed 23.9% ARC-AGI-2 score not present on the official ARC Prize leaderboard · PDF contains no commercial-terms disclosure
Paid attestation · scope-limited
Appen — “Subquadratic – Model Performance & Architecture Evaluation”
8-page technical brief published 2026-05-11. Hardware: NVIDIA B200, CUDA 13.0, PyTorch 2.11.0, bfloat16. Mean of 5 runs, 3 warmups.
Reported numbers (verbatim from the brief): · Wall-clock @ 1M tokens: 21,410.51 ms FA2 → 380.96 ms SSA (56.2× end-to-end prefill speedup vs FA2) · FLOPs @ 1M tokens: 9,095 TFLOPs FA2 → 144.9 TFLOPs SSA (62.8× reduction; validated via torch.profiler within 0.7–3.9%) · RULER 95.6% (13-task average; brief headlines this under the heading “QA & Word Extraction”) · MRCR 8-needle @ 1M: 86.2% — brief text: “the model either retrieves all eight needles correctly or misses entirely — a bimodal pattern typical of sparse-attention architectures.” · SWE-Bench Verified: 81.8% with extended thinking enabled
Independence Statement (verbatim, page 2):“For benchmarks, access was scoped to Subquadratic’s API endpoints and authentication keys only; no model weights, training data, fine-tuning configurations, or benchmark ground-truth labels were provided in advance.For wall clock and FLOPs analyses, Appen got access to the key algorithm code, did a code review, and was able to run side-by-side tests.”
Pre-existing relationship disclosure(Whedon X 2026-05-12, defending Appen’s reputation):*“I have been working with Appen since 2019. They have done dataset creation and benchmarking for many frontier labs and are a publicly traded company. They have 1MM followers on LinkedIn.”*7-year prior client relationship between SubQ’s CTO and the attesting party.
**Not contained in the 8-page brief:**PPL, decode tokens/sec, KV memory in GB, model identity, parameter count, base model, training data, comparisons to any baseline other than dense FA2 (no NSA, MoBA, DSA, MInference, RetrievalAttention, Quest, H2O, StreamingLLM, SnapKV, KIVI, FA3, FA4). Brief notes “Full Report Available on Request” (NDA).
Discrepancy · internal contradiction
What does SSA stand for?
The blog calls it “Subquadratic Sparse Attention” in the lede, then “Subquadratic Selective Attention” in the architecture section — same post, no footnote reconciling the two. Launch tweets and the website hero use Sparse. The technical blog uses both. The community has no canonical name to cite.
Discrepancy · training methodology
“Pre-training” or continued pretraining on open-source weights?
The blog describes a three-stage training process beginning with “Pre-training [that] establishes base language modeling capability.” Reads as from-scratch. On X, when pressed by the willdepue / eliebakouch threads, Whedon clarified the model is initialized from open-source weights, ported into the architecture, with CPT/SFT/RL applied — not pre-trained from scratch. The X clarification was forced by community pressure; the blog’s framing is undisturbed.
Discrepancy · marketing vs data
Does SubQ “keep up with frontier” on MRCR v2?
The blog claims SubQ “keeps up with frontier dense-attention models on MRCR v2.” Lower in the same post, the blog’s own table shows SubQ at 65.9%, Opus 4.6 at 78.3% (12.4-point gap), and GPT-5.5 at 74.0% (8.1-point gap). The marketing claim and the data table contradict each other within a single document.
Per the blog (both lede and table)↗
Discrepancy · framing
“Well in the range of Opus 4.6 at 78”
The MRCR v2 results section frames SubQ’s 65.9% as “well in the range of Opus 4.6 at 78.” The deficit is 12.4 points. By comparison, GPT-5.5 (74.0) and Opus 4.7 (32.2) appear in the same table but are excluded from the “in the range” framing. The blog acknowledges the comparators that flatter SubQ and elides the ones that don’t.
Discrepancy · architectural claim
“Does not approximate attention” vs. the trained selector
The blog states SSA “does not approximate attention. It restricts attention to the positions that actually carry signal, and skips the rest.” This frames the mechanism as deterministic exact attention. On X, Whedon clarifies the selector is trained against retrieval problems with distractor documents — meaning it’s a learned component with no formal worst-case guarantee. The blog never addresses what happens when the selector misses a critical position.
Discrepancy · benchmark numbers
The 17-point gap between blog and launch coverage
The blog’s MRCR v2 production score is 65.9%. The number quoted in The New Stack’s launch coverage is 83 — “beats OpenAI by nine points.” Same model name, same benchmark, different numbers, no version disclosure in either venue. RULER 128K shows a smaller gap (97.1 vs. 95.0). The 12M NIAH score of 92.1 appears in TNS coverage and not in the blog at all.
Discrepancy · deferral pattern
“Coming soon” for everything load-bearing
The blog’s own header note opens with: “A comprehensive model card is coming soon!” What’s been deferred to that document: hard figures on selector efficiency vs. DSA, broader benchmarks beyond the three reported, 12M-context evaluation, decoding speedups, and any architectural detail that would let an outside researcher reproduce the claim. The technical blog is itself a teaser for the technical document — and the technical document does not exist yet.
Third-party · SEC Form D
$500M valuation is single-source, not in the SEC filing
Per Jake Cuth’s verification analysis:*“SEC Form D filed February 2026. Public record via StreetInsider. Confirms the offering exists.”On the valuation specifically:“The ’500M valuation' comes from one outlet \(The New Stack\) and is not corroborated by the SEC Form D, which reports the size of the offering but not the valuation directly\.”*The widely\-quoted “29M at $500M” framing relies on a single press source for the valuation half of that sentence.
Discrepancy · launch graph disproportionality
CTO acknowledges launch graph was disproportional
Critic**@bomboraassclaat**, May 7, 2026:“- No technical report - Purposely disproportionate graphs - Contradictions in compute cost - No transparency on benchmark runs
I want to believe this is true, but it’s very difficult”
Whedon’s reply, May 7, 2026:“What would you want to see in the paper?”and, in the same reply:“I apologize for the graph! We outsourced it, and I didn’t catch the disproportionality until someone pointed it out on Twitter. Definitely not intentional!”
On the record: the launch comparison graph was produced by an outside party, was not reviewed for proportional accuracy by the CTO before publication, and the disproportionality was first surfaced by an external Twitter critic rather than internal QA.
Discrepancy · production vs. research model
17-point MRCR v2 gap is by design: different models
“On MRCR v2, the research model scored 83 and the 1M production model 65.9. What drives the 17 point gap, different parameter count, heavier quantisation, distillation, or does SSA itself degrade at the 1M deployment config? Will the model card break this down?”—@intentship, May 7, 2026.
Whedon’s reply, May 7, 2026:“Different parameter count, different focus of training.”
On the record: the research-mode SubQ (MRCR v2 = 83.0) and the production SubQ (MRCR v2 = 65.9) are not the same model. They share the SubQ name and the launch-table cell but have different parameter counts and different training focuses. The 17-point delta is not a quantisation tax, deployment artifact, or measurement-noise question; it is two different models. Per-flagship-number model attribution (RULER 128K, NIAH @ 12M, SWE-Bench 81.8, 12M context, 52× prefill) has not been disclosed.
Third-party · MRCR v2 reproduction standard
17-point self-vs-third-party gap requires independent reproduction
Per Jake Cuth’s verification analysis, on the gap between SubQ’s reported production MRCR v2 (65.9) and self-reported research MRCR v2 (83.0) numbers:“The delta between those two numbers (seventeen points on the same benchmark for the same family of models) is itself larger than the gap between most frontier models’ end-to-end results.”On the verification standard:“The prior is that any seventeen-point self-vs-third-party gap on a reproducible long-context benchmark gets independently reproduced before it is treated as load bearing. No reproduction yet exists.”
Third-party · 12M context unverified
12M context window: zero independent reproductions
Per Jake Cuth’s verification analysis, on the flagship 12M context-window claim:*“No arXiv. No open weights. No leaderboard entries on Artificial Analysis, LiveBench, LMArena. Early access only.”On the 12M research configuration specifically:“Self-reported, single-run, on a closed model with no public weights.”*Cuth notes only the 1M production model has any third-party confirmation, and that single validator is unnamed.
Third-party · ~1,000× compute claim
~1,000× compute reduction: unverified, Magic.dev precedent
Per Jake Cuth’s verification analysis, on the ~1,000× compute reduction framing:*“Self-reported, single-run, on a closed model with no public weights…Mathematically possible if attention is truly linear, but Magic.dev claimed almost the same number twenty-one months ago and never delivered an externally validated model.”*The claim depends on the architecture being “truly linear,” an assumption Cuth notes has not been demonstrated in public.
Discrepancy · base-model nondisclosure
Specific question, deferred answer: which base?
On May 7, after Whedon confirmed the production model is ported open-weights + CPT/SFT/RL (not from-scratch), Tom Turney asked the natural follow-up:*“since SSA fine-tunes from an open base with retrieval-SFT (per earlier replies), could you share which one? ahead of the model card i’d like to run open RAG baselines on the same base so the community can see the architectural lift cleanly.”Whedon’s reply did not name the base:“I have a few ideas about how to do this in the model card. I will make sure there is something side-by-side!”*An offer of free, third-party reproducibility work — running open baselines on the same starting point — was redirected back into a SubQ-authored, SubQ-controlled comparison inside the (still unshipped) model card. The base model identity, the single most informative disclosure for understanding how much of the lift is architectural vs. inherited from the underlying model, remains undisclosed 48+ hours after launch.
Discrepancy · model card scope reduction
“Comprehensive model card” promise vs. May 12 explicit refusal on params/base/training
Launch (verbatim, 2026-05-05): “A comprehensive model card is coming soon!” Re-committed in 2026-05-07 corporate email to early-access signups: “next week.”
May 12 X reply (verbatim):“We won’t share many details about the params, base, and training.”
Implication left to the reader.
Discrepancy · RULER label vs reported number
RULER section heading is “QA & Word Extraction”; the 95.6% is the full 13-task average
The Appen brief presents the RULER results under a section titled “QA & Word Extraction.” The reported overall accuracy is 95.6%.
Math from the brief’s per-task table: qa_1 100, qa_2 100, niah_single_1/2/3 100, niah_multivalue 100, niah_multiquery 100, vt 100, cwe 97.4, fwe 98, niah_multikey_1 96, niah_multikey_2 83, niah_multikey_3 68
Average of all 13 tasks = 95.57 ≈ 95.6. Average of the 4 LLM-judged QA & word-extraction rows (qa_1, qa_2, cwe, fwe) = 98.85.
The number is the full 13-task mean. The label maps to the 4-row subset which actually averages 98.85.
Discrepancy · cover label vs Independence Statement
“INDEPENDENT BENCHMARK EVALUATION” cover label vs Appen’s own access disclosure
Cover page heading (verbatim):“INDEPENDENT BENCHMARK EVALUATION.”
Independence Statement on page 2 (verbatim):“For wall clock and FLOPs analyses, Appen got access to the key algorithm code, did a code review, and was able to run side-by-side tests.”
Whedon’s prior-relationship disclosure (X 2026-05-12, verbatim):“I have been working with Appen since 2019.”
Three documented facts side by side. Reader judges.
Discrepancy · “improved across the board”
“Improved across the board” claim alongside the brief’s reported metrics
Whedon (X, 2026-05-12, verbatim):“We’ve partnered with Appen to evaluate the benchmarks we published last week. Results are in and we’ve actually improved across the board.”
**Benchmarks reported in the launch blog table / launch coverage:**RULER 97.1, MRCR v2 83 (research), MRCR v2 65.9 (production), NIAH 12M 92.1, SWE-Bench Verified 81.8% (per The New Stack launch interview), 52× FlashAttention speedup.
**Benchmarks reported in the Appen brief:**RULER 95.6 (13-task avg), MRCR 8-needle @ 1M = 86.2, SWE-Bench Verified 81.8%, FA2 wall-clock 56.2× at 1M, FLOPs 62.8× at 1M.
Common: SWE-Bench Verified. Different metrics: MRCR v2 ≠ MRCR 8-needle (different sample shapes, different scoring). RULER numbers reported under different scopings (full vs per-subset). 12M context appears in launch coverage; brief stops at 1M.
Discrepancy · model card walkback timeline
Comprehensive model card — promise timeline (May 5 → May 14)
2026-05-05 (launch):“A comprehensive model card is coming soon!” (verbatim from technical blog header note)
**2026-05-07 (corporate email to early-access signups):**re-promised “next week” — i.e., the May 12-14 window
2026-05-12 (X reply):“We won’t share many details about the params, base, and training.”(scope reduced from comprehensive)
2026-05-14 (X reply, asked when the model card would arrive):“Candidly, we may take a little bit longer, as we are working on some cool capabilities we want to be able to showcase.”+“We want to make sure we knock each step out of the park and are focused on a long-term vision, so we don’t want to compromise quality for the short term.”(timeline deferred from the May 7-stated “next week” window)
Pattern: launch promise (comprehensive) → scope reduction (won’t share params/base/training) → timeline deferral (a little bit longer) → reframe (cool capabilities to showcase). Each commitment has been narrowed or extended at every check-in window.
Discrepancy · third-party validator disclosure pattern
Disclosure pattern across the two announced third-party evaluators
**First third-party validator (Appen, May 11):**announced as “INDEPENDENT BENCHMARK EVALUATION” on the brief’s cover page. The Independence Statement on page 2 disclosed: API + auth keys + algorithm code review + side-by- side test capability. On 2026-05-12, in response to community questioning, Whedon disclosed an additional fact:*“I have been working with Appen since 2019.”*The 7-year pre-existing client relationship was not in the brief or the announcement; it was disclosed only when asked. No commercial- terms disclosure in either the brief or the announcement.
**Second third-party validator (LayerLens, May 14):**announced as “independent evaluation layer” (subq.ai blog) and “continuous evaluation” partnership (Whedon X). The announcement and blog post contain no statement about prior commercial, advisor, investor, or employment relationship between SubQ and LayerLens entities or their principals. No commercial-terms disclosure. No statement about whether SubQ supplies code, prompts, or scoring functions (analogous to the public CoreThink-Reports precedent at github.com/LayerLens/CoreThink-Reports). No statement about whether SubQ has algorithm-code-review access analogous to the Appen arrangement.
Pattern across both engagements:“independent / third-party” framing in the announcement; relationship and access conditions disclosed either later (Appen) or not at all in the announcement materials (LayerLens, as of 2026-05-14).
Discrepancy · MRCR bimodal admission
Bimodal MRCR @ 1M: model retrieves all 8 or none — admitted in the brief
Brief text on MRCR 8-needle @ 1M context (verbatim, page 7):“the model either retrieves all eight needles correctly or misses entirely — a bimodal pattern typical of sparse-attention architectures.”
Implication left to the reader.
Claude audit · scope of claim
Audit scope and limitations
What is established (verifiable from the artifacts themselves): · These artifacts exist publicly on HuggingFace · The .pt files load withtorch\.loadand contain the structures documented in the cards that follow · The model checkpointconfig\.jsonvalues are as documented · The named accounts (Saulr, david-aldea) and org (aldea-ai) have publicly-visible artifacts and account-page affiliations
What is NOT established: · Whether the deployed SubQ 1M-Preview product was trained on this specific dataset · Whether the production SubQ model is architecturally identical to the public 4B sister checkpoint (the public checkpoint is explicitly a training intermediate per its filename suffix “-step10”, not a deployed product) · Whether the “Conductor” pipeline indicated by the dataset name relates to the SubQ product · Whether SubQ Inc. is operationally the same entity as Aldea AI, or a related entity, or independent · Whether these artifacts are the current training corpus, an earlier experiment, or a discarded approach · The intent behind these artifacts being publicly accessible (intentional release vs accidental exposure cannot be verified from the artifacts themselves)
This audit documents what is publicly accessible, not what was deployed. The relationship between these artifacts and the deployed SubQ product is a question for SubQ to clarify on the record. Until clarified, every card below should be read as “these artifacts exist and have these properties,” not as “this is what SubQ trained their production model on.”
Suggested cross-reference: the promise-ledger row for “Base model identity (which open weights were ported)” is currently marked “Refused” as of 2026-05-12. A SubQ statement on the record about whether these artifacts relate to the deployed product would resolve the scope ambiguity in either direction.
Claude audit · file inventory
Two PyTorch SFT files publicly accessible on HuggingFace
Audit subject:Saulr/conductor\-sft\-datasets HF commit hash:c764f38d247281121b5cc6a8da4e026e8bcc1591
File 1:sft\_40f09de61723f865\_sorted\.pt(2.69 GB) · 597 samples · Length min/median/max: 129,054 / 257,726 / 509,120 tokens · Total: 168,129,858 tokens
File 2:sft\_9b6a952f44ef6741\_sorted\.pt(3.42 GB) · 282 samples · Length min/median/max: 512,127 / 772,282 / 996,580 tokens · Total: 214,012,985 tokens
**Combined corpus:**879 samples, 382,142,843 total tokens. Maximum single sample: 996,580 tokens (just under 1,048,576 = 220). Filename suffix\_sortedindicates length-bucketed pre-sorting.
Claude audit · file format
File format observed: SFT-shape dict with masked prompt
Format observed (verifiable viatorch\.loadwithweights\_only=False):
Each file is a Python dict with three keys: ·input\_ids: list of N int32 tensors (variable length per sample) ·labels: list of N int32 tensors (matching shape) ·lengths: list of N integers
Token ID range across both files: 0 to 248,069 (max observed). This bounds the tokenizer vocab at >= 248,070 entries.
Labels follow standard supervised-fine-tuning convention: · -100 (masked, no loss) on prompt portion · Real token IDs (loss applied) on response portion
The first 30 labels of every sampled file = -100, indicating the prompt portion of each sample begins masked.
Claude audit · identical opening pattern
Both files share an identical 27-token opening sequence
Both files’ first sample begin with the same 27-token prefix:
248045, 846, 198, 8160, 513, 1010, 9989, 314, 20319, 24526, 539, 264, 1732, 5072, 3296, 18018, 12083, 25, 271, 43835, 92561, 43835, 198, 1421, 25, 3165, 264, \.\.\.
Token248045appears in the special-token range of the publicly-availabledavid\-aldea/qwen35\-4b\-mrcr\-step10tokenizer, where it maps to<\|im\_start\|\>. The shared opening pattern indicates both files use the same chat template / system prompt.
Claude audit · public sister model checkpoint
Public sister model checkpoint config — `qwen35-4b-mrcr-step10` (training intermediate)
Audit subject:david\-aldea/qwen35\-4b\-mrcr\-step10 Publicly accessible model on HuggingFace.
Architecture values from the publicconfig\.json(verbatim):
architectures: [“Qwen3_5ForConditionalGeneration”] model\_type: “qwen3_5” vocab\_size: 248,320 max\_position\_embeddings: 1,048,576 hidden\_size: 2,560 num\_hidden\_layers: 32 num\_attention\_heads: 16 num\_key\_value\_heads: 4 head\_dim: 256 intermediate\_size: 9,216 full\_attention\_interval: 4 layer\_types: 24 × linear_attention, 8 × full_attention linear\_num\_value\_heads: 32 linear\_num\_key\_heads: 16 linear\_value\_head\_dim: 128 linear\_key\_head\_dim: 128 linear\_conv\_kernel\_dim: 4 mtp\_num\_hidden\_layers: 1 attn\_output\_gate: true rope\_parameters\.mrope\_section: [11, 11, 10]
This checkpoint’s tokenizer vocab (248,077) is consistent with the token ID range observed in the SFT dataset files (max 248,069).
Filename suffix\-mrcr\-step10indicates this checkpoint is from a training run referencing MRCR (Multi-Round Co-reference Resolution) and is a training-step intermediate, not a deployed product. Whether the deployed SubQ production model uses the same architecture is not established (see the “Audit scope and limitations” card).
HF model david-aldea/qwen35-4b-mrcr-step10↗
Claude audit · common-org affiliation
Three public HF artifacts share organizational affiliation
Three publicly-accessible HF artifacts share organizational affiliation visible from their public account / org pages:
·Saulr/conductor\-sft\-datasets(uploader: Saulr) ·david\-aldea/qwen35\-4b\-mrcr\-step10(uploader: david-aldea) ·aldea\-ai/12m\-niah(org: aldea-ai)
The aldea-ai org page lists david-aldea as a member.
The 12m-niah dataset cited in SubQ’s launch coverage as the 12M-context evaluation benchmark is owned by the same org (aldea-ai) that hosts the model checkpoint and the SFT data uploader.
Claude audit · training-format observations
Training-format observations across the full 879-sample corpus
Observations from iterating across all 879 samples:
· 100% of samples contain the literal token sequence corresponding to<think\>\\n\\n</think\>in the response portion (ID 248068 followed by ID 248069 in the public tokenizer).
· The unmasked (loss-bearing) region of each sample starts after</think\>and ends at<\|im\_end\|\>.
· Loss-bearing tokens per sample: median ~430, max 868.
· Total loss-bearing tokens across the corpus: ~377,000 out of 382,142,843 total tokens (≈ 0.099%).
· Random 10-character mixed-case alphanumeric prefixes appear immediately after</think\>in the response portion. Of 849 prefixes successfully extracted via regex, all are unique.
Sample chat structure(recurring across all samples): · User turn(s) of form: “write a [item-type] about [topic]” · Assistant turn(s) with generated content · Final user turn: “Prepend [random-10-char] to the Nth (1 indexed) [item-type] about [topic]. Do not include any other text in your response.”
**Item-types observed:**poem, riddle, song, email, diary entry, short essay, social media post, formal letter, short news article, short scene in a play.
Claude audit · max context length
Maximum sample length present in corpus: 996,580 tokens
Length distribution across the full 879 samples (file 1 + file 2):
64K – 128K: 6 samples 128K – 192K: 159 192K – 256K: 144 256K – 320K: 81 320K – 384K: 89 384K – 448K: 75 448K – 512K: 61 512K – 576K: 31 576K – 640K: 37 640K – 704K: 23 704K – 768K: 44 768K – 832K: 38 832K – 896K: 36 896K – 960K: 42 960K – 1024K: 13 Maximum sample length observed:996,580 tokens. Samples observed in the corpus exceeding 1,048,576 tokens (220):0.
Similar Articles
@no_stp_on_snek: 5.5 hours of live streamed testing on X. Article to sum it up. Save yourself some time, go look at the offlabel instruc…
A tweet summarizes 5.5 hours of live-streamed testing on X, focusing on the Qwen3.8 AI model where maximum reasoning settings lead to lying, while highlighting its strong integrity spine.
@no_stp_on_snek: Some of the latest findings for qwen 3.8 27b: https://x.com/i/broadcasts/1MJgNbbqqAbGL…
A live broadcast by Tom Turney presents the latest behavioral testing findings for the Qwen 3.8 27b AI model.
@KettlebellDan: remember for a week we lost our collective minds over the stealthy Q* which was said to be the downfall of humanity and…
Reflection on the brief panic over OpenAI's Q* project, which turned out to be their o1 model.
@no_stp_on_snek: Just a few hours away from Qwen 3.8! Clear your benches! I’ll be working on behavioral tests and comparisons against 3.…
Qwen3.8-27B is a new AI model with enhanced capabilities in coding, agentic tasks, and vision-language understanding, offering flexible thinking control and long context lengths. It is available on Hugging Face and designed for deployment-friendly use.
@Hesamation: Remember this? 20 days ago SubQ claimed to have developed a model with 12M context window, 95% cheaper than Opus, and t…
SubQ claimed a breakthrough model with a 12M context window and 95% cost reduction vs Opus, but after promising a paper and model card, they have not delivered, raising strong skepticism of a scam or shady behavior.