Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Summary
This paper introduces Chinese-Jev, a System One model designed to improve decision-making accuracy and efficiency for Chinese-language tasks, with domain-specific fine-tuning and a new benchmark, CJ-Bench, for evaluation.
View Cached Full Text
Cached at: 09/30/26, 08:22 AM
Paper page - Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Source: https://huggingface.co/papers/2609.36965
Abstract
SystemOnemodelssuchasJevofferanefficientalternativetogenerativelanguagemodelsfortasksthatrequiredecisionsratherthanopen-endedresponses.However,existingJevmodelsexhibitlimitedChinese-languagedecisionaccuracy,restrictingtheirutilityinbothgeneralandspecializedsettings.Inthispaper,weintroduceChinese-Jev,aSystemOnemodelthataddressesthisgapthroughaunifieddataprocessingandtrainingpipeline.OurdataprocessingprotocolconvertsheterogeneousChinese-languageannotationsintoprobabilitytargetsovercandidateoptions,enablingasharedtrainingformulationacrossdomainsandquestionformats.Toenableefficientinference,Chinese-Jevadoptsalightweightencoder-onlybackbonefortextencodingandlearnstoscorecandidateanswersthroughdecision-orientedtraining.Toaddressthemisalignmentbetweenthepre-trainingdistributionanddownstreamChinese-languagescenarios,wefirsttrainthemodelonageneral-purposecorpusof10millionexamples,thenfine-tuneitseparatelyforthemedical,legal,andfinancialdomains.Toevaluatedecisionaccuracyandcalibrationinbothgeneralanddomain-specificChinese-languagesettings,weintroduceChinese-JevBench(CJ-Bench).Afterfirst-stagepre-training,Chinese-Jevexceedstheaccuracyoftheclosed-sourceJevmodelby1.24%ongeneral-domaintaskswhileachievinga20.3xspeedup.Subsequentdomain-specificfine-tuningyieldsa4.0%accuracyimprovementoverJevinmedicineandachieves92%ofJev’saverageaccuracyacrossspecializeddomains,witha17xspeedupandanaveragelatencyofonly15msperexample.Wefurtherdemonstrateon-devicedeploymentofanINT8-quantizedmodelonmobiledevices,achievinganinferencelatencyofapproximately1.0secondperdecision.Theprojectisavailableathttps://gulucaptain.github.io/Chinese-Jev/.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.36965
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.36965 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.36965 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.36965 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@LangChain: We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could o…
This article evaluates Jev, a System One model from TypeSafe AI, as a new agent evaluator, showing it outperforms LLM judges in consistency, speed, and cost.
@svpino: Jev is incredibly good! If you haven't heard, Jev is a new "System One" model optimized for decision-making. For exampl…
Jev is a new AI model optimized for fast and cheap decision-making in classification tasks, demonstrated with a hotel review classifier project using Apify and GPT-5-mini.
Show HN: JevBench, a reproducible benchmark for typed decision models
JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.
Jev introduces a new shape of LLM - System One, aka Decision Models
Jev is a new AI model introduced by TypeSafe AI, categorized as a System One or Decision Model, that outputs floating-point numbers for classification tasks like spam detection and ranking with competitive pricing.
Been experimenting with Jev — interesting approach to AI agents
The author discusses experimenting with Jev, a tool for AI agents focused on decision-making, which claims significant speed and cost benefits compared to using large language models for all tasks.