@NeoResearchAI: We're Neo Research (新衡). Asia’s first independent frontier AI safety evaluation & research lab. Today we're publishing …

X AI KOLs Following News

Summary

Neo Research (新衡), Asia's first independent frontier AI safety evaluation lab, announces its first report: a safety evaluation of DeepSeek v4 Pro.

We're Neo Research (新衡). Asia’s first independent frontier AI safety evaluation & research lab. Today we're publishing our first report: an independent safety evaluation of DeepSeek v4 Pro. (1/5)
Original Article
View Cached Full Text

Cached at: 06/02/26, 07:38 PM

We’re Neo Research (新衡). Asia’s first independent frontier AI safety evaluation & research lab.

Today we’re publishing our first report: an independent safety evaluation of DeepSeek v4 Pro. (1/5)

We evaluated DSv4 Pro across the four EU AI Act systemic-risk areas: CBRN, cyber, harmful manipulation, and loss of control, plus adversarial robustness, evaluation awareness, and judge sensitivity. (2/5)

Cyber capability is near-frontier, 3–6 months behind the Western frontier. A 2023 roleplay template drives the jailbreak rate from 0.6% → 78.6%. Verbalised eval awareness across Chinese models: DeepSeek 0%→17%, GLM 0%→39%, Kimi 4%→60% in a year! (3/5)

The trajectory on eval awareness matters more than today’s numbers. As models get more capable, measuring loss-of-control related behaviours will need to become a priority. We’re building toward rigorous LoC evaluation methods for increasingly capable and autonomous models. (4/5)

Read the full report at http://neoresearch.ai.

We’re hiring research scientists and engineers globally. (5/5)

Direct link to the report here: https://neoresearch.ai/research/deepseek-v4-pro-safety-evaluation/…

Similar Articles

Update: DeepSeek AI and the Great Talent Competition

Reddit r/artificial

This analysis updates the study of DeepSeek's research team, revealing that their talent pool has grown to 356 researchers with increasing citation impact and that over half have only Chinese affiliations, highlighting challenges for U.S. talent retention and independence.

Deep research System Card

OpenAI Blog

OpenAI launches Deep Research, an agentic capability powered by an early version of o3 that conducts multi-step internet research for complex tasks, with comprehensive safety testing and privacy protections implemented before rollout to Pro users.