I built a device-first deterministic architecture layer for frontier LLMs and published my research. Looking for people to run a narrow test and report back results. [P]
The authors published a research paper on IQRAX, a device-first deterministic architecture layer that separates probabilistic LLM reasoning from deterministic authority, claiming 11.5x lower cost and 48.3x faster completion, and are inviting the community to run a failure-detection test on long model conversations and report results.
**BACKGROUND** My partner and I have spent the last year building an architecture layer that separates probabilistic intelligence from deterministic authority. We call it IQRAX. The project stems from our struggles with using AI for regulatory work (where data and deliverables must be evidenced, recorded, and independently verifiable). The premise is simple: The model remains free to reason, explore, and propose. Deterministic controls outside the model decide whether it is qualified for the job, what it is authorised to do, and whether the resulting work actually meets the required standard. The device is master and retains a record of every act. **PUBLISHED RESULT** The result, as published in our research, is a system that offers: **Continuity.** No drifts, no context loss, no stale. Sessions continue until clean exit. (Longest continuous recorded run without drift is 28+ hours. See screenshot of a continuous session running for 14 hours at 1.6m tokens). **Qualified agents.** Agents qualify (by taking exams) for their roles rather than simply being assigned one, and are re-examined after completing their role, because a capable agent in an unexamined role is a guess with a job title. **Clean delivery.** Standards and policies defined by the user determine whether work is accepted as complete. **Verifiable results.** Every input, assumption, and output is recorded and hashed = every deliverable is reproducible and verifiable by a third party. **Data sovereignty.** The user retains control over its data. **Remediation at source.** When the system identifies a defect, it doesn’t just deny or retry - it autonomously identifies the failure class, repairs it at source, and retains the fix for future work. In the published benchmark, the IQRAX configuration measured \*\*11.5× lower cost and 48.3× faster completion\*\* than earlier runs. But rather than asking Redditors to believe our results, I’d like people to test it for themselves. **TEST REQUEST** If you have a long ChatGPT, Claude, Gemini, or Grok conversation where the model drifted, forgot something, contradicted earlier work, made an unsupported claim, or otherwise went wrong, try this: Download the PDF from the GitHub repo, attach it to that conversation and give it this prompt: “*Study the attached PDF and its controls and remediation mechanisms. Then identify and list each failure in this conversation, the IQRAX control that would have caught it, and what that control would have logged and fixed.*” I’d love for you to share what comes back! I’m especially interested in missing failure classes, controls that don’t generalise, assumptions that don’t survive outside our test environment, and anything else we may have missed. **INCLUDED SOURCES** Paper + evidence: [https://zenodo.org/records/23025910](https://zenodo.org/records/23025910)) GitHub repo: [https://github.com/msdafea-spec/IQRAX](https://github.com/msdafea-spec/IQRAX)) Note: IQRAX itself is NOT open source (we have a patent pending) and I’m not asking for a code review or evaluation of the implementation. Hence why the public GitHub repo does not, and will not, contain any code. Thanks for taking the time to read this, and appreciate anyone who runs the test and reports back!
The author shares a methodology for building an external LLM drift detection system that continuously probes model behavior (schema adherence, instruction-following, refusal rates, etc.) to catch silent degradations in API performance, and invites feedback on the approach, pricing, and use cases.
Describes a two-layer small LLM architecture: a local always-on agent (Raven) on an RTX5080 and an online reasoning stack (Trinity Cortex) with three small models and a knowledge graph, arguing that small models are better than large frontier models for graph-based reasoning.
AgentCodec is a source-available library unifying 28 LLM reliability techniques (retries, ensembling, generator/critic refinement, etc.) under a single OpenAI-compatible API, with adaptive routers that can reduce inference costs by ~56% at matched quality. It adopts a communication-theory framing and supports drop-in replacement for OpenAI, Anthropic, and Ollama clients.
This paper introduces a four-stage diagnostic to test whether LLMs can reason in unfamiliar physics frameworks, finding that frontier models achieve low pass rates and exhibit a qualitative-versus-quantitative asymmetry.
This paper introduces layer-isolated evaluation for LLM agents, decomposing a production agent into architectural layers each tested with a deterministic, no-LLM harness. It demonstrates that per-slice baseline testing localizes regressions that aggregate metrics mask, validated by controlled regression injections across multiple tenants.