Tag
The paper introduces a curvature-based cryptanalysis method to extract hidden feed-forward network structures in transformers using black-box queries, achieving high-fidelity functional model substitutes.
This paper proposes GraphRP, a proactive defense framework using model reprogramming to protect GNNs from model extraction attacks, with a structure-aware gating mechanism that preserves benign utility while degrading adversarial queries.
A research paper introducing a multi-stage forward-query method to cryptanalytically extract isolated bias-free GLU feed-forward block weights, demonstrating sub-percent recovery accuracy on Qwen, Llama, and Gemma components while noting end-to-end model API attacks remain unsolved.
This paper introduces ADS-C, an antidistillation defense for classification that provably preserves top-1 accuracy while degrading student model performance by up to 29.7 percentage points, achieving zero utility cost for the teacher.
This paper introduces Reasoning Exposure Prompting (REP), a method that uses shadow-model demonstrations in code-like formats to elicit hidden reasoning traces from LLMs, showing that interface-level trace hiding is insufficient to prevent extraction of useful reasoning signals.
This paper presents the first model extraction attack on graph classification under strict black-box constraints, exploiting subgraph explanations to estimate decision boundaries. The findings reveal that mandated explainability interfaces create exploitable security vulnerabilities in Graph Neural Network services.
A user reports that the 3.6 GB Gemma 4 e4b model extracted from Google AI Edge Gallery on Android outperforms larger 3.7 GB Unsloth versions and community ports, raising questions about hidden optimizations.