I Don't Want to Interact with Stochastic Parrots
Summary
The article expresses a reluctance to interact with AI systems referred to as 'stochastic parrots,' likely critiquing the limitations and ethical concerns of large language models.
Similar Articles
Evaluation Methodologies (Evals) in AI Engineering
Course notes for Chapter 3 of AI Engineering, covering evaluation methodologies (Evals) in depth: starting from the challenges posed by open-ended outputs and black-box models, it introduces language modeling metrics such as cross-entropy, perplexity, BPC/BPB, and compares two mainstream evaluation approaches — AI-as-a-judge and comparative evaluation.
Has anyone noticed this trend toward writing/speaking style among newer models (both open and closed models). They are trending toward information density and expanded vocabulary. It's not quite 'caveman speak' but trending that way.
An observer notes that newer LLMs, both open and closed, are converging on an information-dense, tersely worded writing style with expanded, sometimes esoteric vocabulary — compared to William Gibson — which may come at the cost of readability for average users.
FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
This paper introduces FAER, an auditable full-trajectory replay framework for language model post-training that formalizes the gap between cache-level selection feedback and downstream learner utility, showing that a learner-aware selector (FAER-UTILITY) outperforms format-feedback and uniform baselines on GSM8K with Qwen2.5-1.5B-Instruct.
A First Glance at Jev for Network Traffic Classification: Accuracy, Processing Time, and Cost
This paper presents the first empirical study evaluating Jev, a general-purpose decision model, for network traffic application classification using only early-flow packet features on the CESNET-QUICEXT-25 dataset. Labeled examples boost Jev's accuracy from 9.80% to 34.50%, but trained tree ensembles and GPT-5.6 Sol still outperform it, suggesting labeled in-context examples alone are insufficient to match dedicated classifiers.
Large Language Bayes Is Not Reparameterisation-Invariant
This paper shows that Large Language Bayes (LLB), which averages posteriors from language-model-generated probabilistic programs using weights from an exponentiated evidence bound, is not reparameterisation-invariant. Equivalent program writings can receive substantially different weights (up to 31.9× discrepancy that can even reverse sign), inverting Bayes factors and inducing error in model posteriors controlled by the spread of bound shortfalls via a sharp Hilbert-distance bound.