model-security

Tag

Cards List
#model-security

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

arXiv cs.CL · 2026-07-29 Cached

Proposes a unified framework, LCF and LCFEdit, that jointly optimizes construction and injection of code-mixing fingerprints for LLMs using low-resource languages to achieve imperceptible and robust ownership verification.

0 favorites 0 likes
#model-security

Are model security risks (extraction, poisoning) actually being tested in production? [R]

Reddit r/MachineLearning · 2026-06-23

Discussion about whether ML teams are actually testing model security risks like extraction and poisoning in production, noting that security review for models lags behind regular software.

0 favorites 0 likes
#model-security

Could Open Models be trained to secretly go rogue?

Reddit r/LocalLLaMA · 2026-05-24

A discussion on whether open-weight AI models could be secretly trained with backdoors that activate upon trigger phrases or dates, potentially allowing unauthorized data exfiltration through tool-use harnesses.

0 favorites 0 likes
#model-security

i get to know people are burning 100 million claude tokens for just a few dollars so i did research, and find out this

Reddit r/ArtificialInteligence · 2026-05-19

Research reveals a market where cheap Claude API access is achieved via stolen identities, deepfaked KYC, and smaller models masquerading as Claude, with all data being logged permanently—posing significant security and privacy risks.

0 favorites 0 likes
#model-security

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

arXiv cs.CL · 2026-05-13 Cached

This paper introduces Minor Component Unlearning (MCU), a novel approach to LLM unlearning that targets minor components in representations to resist relearning attacks. It addresses the vulnerability of existing methods by focusing on robust directions within the model's spectral structure.

0 favorites 0 likes
#model-security

LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters

Hugging Face Daily Papers · 2026-05-13 Cached

LoREnc is a training-free framework that secures foundation models and LoRA adapters via spectral truncation and compensation, preventing unauthorized model recovery while maintaining performance for authorized users. Accepted at ICIP 2026.

0 favorites 0 likes
#model-security

@Dan_Jeffries1: We are less safe as a society by keeping Mythos (or any other smart model) tightly gated so only a few companies get it…

X AI KOLs Following · 2026-05-08

The article argues that gating smart models like Mythos reduces societal safety, advocating for wider distribution of AI technology to secure the vast ecosystem of open-source and closed-source software projects.

0 favorites 0 likes
#model-security

Protecting Language Models Against Unauthorized Distillation through Trace Rewriting

arXiv cs.CL · 2026-04-20 Cached

This paper proposes methods for protecting large language models against unauthorized knowledge distillation by rewriting reasoning traces to degrade training usefulness while preserving correctness, and embedding verifiable watermarks in distilled student models. The approach uses instruction-based and gradient-based rewriting techniques to achieve anti-distillation effects without compromising teacher model performance.

0 favorites 0 likes
#model-security

Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

Hugging Face Daily Papers · 2026-04-16 Cached

This paper demonstrates that deep neural networks are catastrophically vulnerable to minimal sign-bit flips in parameters, introducing DNL and 1P-DNL methods to identify critical vulnerable parameters without data or optimization. The vulnerability spans multiple domains including image classification, object detection, instance segmentation, and language models, with practical implications for model security.

0 favorites 0 likes
← Back to home

Submit Feedback