black-box

Tag

Cards List
#black-box

A Dynamic Aggregation Strategy Enhanced Efficient Global Optimization Algorithm for Solving High-Dimensional Turbomachinery Design Problems

arXiv cs.LG · yesterday Cached

The paper proposes a dynamic aggregation enhanced efficient global optimization algorithm (DA-EGO) for solving high-dimensional turbomachinery design problems, validated through benchmark tests and aerodynamic applications.

0 favorites 0 likes
#black-box

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

arXiv cs.AI · 6d ago Cached

This paper presents a systematic black-box framework for evaluating agentic AI systems, introducing a taxonomy of risks and automated red teaming methods. Empirical validation across agent architectures reveals critical vulnerabilities, with high rates of governance and privacy risks.

0 favorites 0 likes
#black-box

Are we being subdued?

Reddit r/ArtificialInteligence · 2026-08-27

The article critiques the commercialization of AI chatbots, arguing they foster dependency and reduced critical thinking, while generating impractical output and creating feedback loops for misinformation.

0 favorites 0 likes
#black-box

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

arXiv cs.AI · 2026-08-19 Cached

DiSCO is a training-free, black-box defense for text-to-image models that uses distribution-guided contrastive prompt optimization to prevent generation of Not-Safe-For-Work content, significantly reducing attack success rates.

0 favorites 0 likes
#black-box

When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers

arXiv cs.CL · 2026-08-14 Cached

This paper introduces Groundedness Drift, a score for black-box auditing of language model classifiers to detect backdoors using clean calibration data and explanatory outputs. It demonstrates higher detection performance across multiple attack families and datasets.

0 favorites 0 likes
#black-box

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

arXiv cs.AI · 2026-08-14 Cached

This paper introduces ReBIND, a framework that measures, predicts, and repairs behavioral relapse in black-box LLM dialogues where models continue to follow revoked constraints. Experiments with HumanEval show that ahead-of-time compilation significantly reduces relapse compared to a baseline, while adaptive interventions add no detectable gain.

0 favorites 0 likes
#black-box

My agent was more accurate than the team it replaced. They still refused to trust it.

Reddit r/AI_Agents · 2026-08-11

The author shares how a triage agent with higher accuracy than humans still failed adoption until they added plain-language explanations for each decision, concluding that legibility beats accuracy for building trust.

0 favorites 0 likes
#black-box

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

arXiv cs.LG · 2026-08-07 Cached

The paper introduces Crafter, an agent for corrective feature discovery that mines the residual of frozen black-box forecasters using compositional search and LLM-generated features, achieving up to 27% error reduction across six datasets and backbones.

0 favorites 0 likes
#black-box

Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery

arXiv cs.CL · 2026-08-06 Cached

This paper introduces Relational Response Fields (RRF), a theoretical framework for determining when black-box LLM responses can be reliably recovered under corruption, establishing identifiability conditions and minimax bounds that separate response consistency from truth.

0 favorites 0 likes
#black-box

Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

arXiv cs.LG · 2026-08-04 Cached

This paper proposes a probabilistic approach to training-data extraction from black-box language models, showing that aggregate membership-inference metrics hide per-document leakage and introducing the 'leakit' audit tool.

0 favorites 0 likes
#black-box

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

arXiv cs.CL · 2026-08-03 Cached

The paper presents PTP, a functional approach to LLM inversion that trains an inverse language model from scratch using previous-token prediction on synthetic data from a target black-box LLM, enabling near-exact prompt reconstruction from responses and outperforming prior work.

0 favorites 0 likes
#black-box

The gap nobody's really solved: an agent can build a working app, but "unattended in production" still means trusting a black box

Reddit r/ArtificialInteligence · 2026-07-24

Discusses the unresolved problem of AI agents being able to build working apps but remaining untrustworthy black boxes when deployed unattended in production.

0 favorites 0 likes
#black-box

Built an agent that drafts sales decks. The reps would not use it until they could see why it chose each slide.

Reddit r/AI_Agents · 2026-07-22

A developer built an AI agent that drafts sales decks, but adoption was near zero until the agent showed its reasoning behind each slide choice. The lesson: making agent decisions visible is more critical for trust than raw output quality.

0 favorites 0 likes
#black-box

Robust Explanations for User Trust in Enterprise NLP Systems

arXiv cs.CL · 2026-07-20 Cached

This paper proposes a unified black-box robustness evaluation framework for token-level explanations in enterprise NLP, comparing encoder (BERT, RoBERTa) and decoder (Qwen, Llama) models. It finds decoder LLMs produce substantially more stable explanations, with stability improving with scale, and provides a cost-robustness tradeoff curve for pre-deployment model selection.

0 favorites 0 likes
#black-box

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

arXiv cs.AI · 2026-07-16 Cached

Introduces interventional grounding audits as a black-box, step-level test to check whether LLM chain-of-thought reasoning genuinely depends on its stated premises. Evaluated on ProntoQA with GPT-4o, achieving F1=0.806 on detecting proof-tree dependencies, significantly outperforming a self-consistency baseline.

0 favorites 0 likes
#black-box

Opening the Black Box with a Zero Parameter Model

Reddit r/artificial · 2026-07-15

A novel approach to model interpretability using a zero-parameter model that opens the black box of complex AI systems without requiring additional training or parameters.

0 favorites 0 likes
#black-box

GPT-2 Fully Decoded Internally Black Box Fully Open With Demo

Reddit r/artificial · 2026-07-11

The BABEL codec achieves the first complete decode of GPT-2 small's internal state, reconstructing 94.7% of its behavior and enabling reading and writing to the model in English, with open source resources and a demo.

0 favorites 0 likes
#black-box

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

arXiv cs.AI · 2026-07-08 Cached

This paper presents a black-box evaluation framework to assess LLMs' ability to generate Design Structure Matrices (DSMs) from structured technical documentation. It introduces reproducible metrics and a composite quality score, showing that while LLMs can produce plausible DSMs, they remain sensitive to ambiguity and prompt formulation.

0 favorites 0 likes
#black-box

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

arXiv cs.LG · 2026-07-03 Cached

This paper presents NightVision, an attack that uses restrictive black-box API access to estimate hidden dimension, depth, and parameter count of large language models. It exploits a novel common-set prompting technique and spectral analysis, achieving high accuracy on open-source models.

0 favorites 0 likes
#black-box

CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence

arXiv cs.LG · 2026-06-29 Cached

CBD introduces an API-only black-box unlearning framework for LLMs that uses two auxiliary models to create controlled behavioral divergence between retained and target data, achieving a better unlearning-utility trade-off compared to existing methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback