初探 Jev 在网络流量分类中的表现:准确率、处理时间与成本
摘要
本文首次对 Jev(一种通用决策模型)进行实证评估,研究其仅利用早期流包特征在 CESNET-QUICEXT-25 数据集上进行网络流量应用分类的表现。带标签的示例可将 Jev 的准确率从 9.80% 提升至 34.50%,但经过训练的树集成模型和 GPT-5.6 Sol 仍然优于它,这表明仅靠带标签的上下文示例不足以媲美专用分类器。
arXiv:2610.00376v1 Announce Type: new
Abstract: We evaluate Jev on ten dataset-defined application labels in CESNET-QUICEXT-25 using only the first ten packets' sizes, directions, and inter-packet times. To the best of our knowledge, this is the first empirical study of general-purpose decision models, represented here by Jev, for application classification of network flows. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev's accuracy from 9.80% to 28.42%. Random Forest and Extra Trees trained on 8,000 records achieve 69.95% and 66.80% and outperform Jev in every week. Increasing Jev's context to 150 examples yields 34.50% on the first test week. On a paired 100-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0.750 s, versus 37% and 6.036 s for the generative language model OpenAI GPT-5.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause.
查看缓存全文
缓存时间: 2026/10/03 09:52
# A First Glance at Jev for Network Traffic Classification:Accuracy, Processing Time, and Cost Source: [https://arxiv.org/html/2610.00376](https://arxiv.org/html/2610.00376) Shenghe Xu††thanks:Work performed outside the author’s role at Amazon\. Email:[shenghexu@gmail\.com](mailto:[email protected])\.Lifan Mei††thanks:Corresponding author:[lifan\.mei@xjtlu\.edu\.cn](mailto:[email protected])\.Affiliation:Xi’an Jiaotong\-Liverpool University September 30, 2026 ###### Abstract We evaluate Jev on ten dataset\-defined application labels in CESNET\-QUICEXT\-25 using only the first ten packets’ sizes, directions, and inter\-packet times\. To the best of our knowledge, this is the first empirical study of general\-purpose decision models, represented here by Jev, for application classification of network flows\. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev’s accuracy from 9\.80% to 28\.42%\. Random Forest and Extra Trees trained on 8,000 records achieve 69\.95% and 66\.80% and outperform Jev in every week\. Increasing Jev’s context to 150 examples yields 34\.50% on the first test week\. On a paired 100\-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0\.750 s, versus 37% and 6\.036 s for the generative language model OpenAI GPT\-5\.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges\. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations\. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause\. ## 1Introduction The appeal of general\-purpose language models, made familiar through ChatGPT, extends beyond producing text\. Instructions and a few examples can describe a task to an existing model without updating its weights\[[1](https://arxiv.org/html/2610.00376#bib.bib14)\]\. For networking, this suggests a practical possibility: using a shared model service to interpret measurements and supply decisions, reducing the need to build a separate predictor for each task\. Research on adapting language models to networking has begun to explore this possibility\[[2](https://arxiv.org/html/2610.00376#bib.bib15)\]\. Whether that flexibility translates into useful decisions depends on the observations available, the supervision supplied, and the operational cost of each prediction\. Early traffic classification brings these requirements together\. Network monitoring and service management need to associate flows with applications, while encryption limits directly readable content and waiting for a complete flow delays identification\. Packet sizes, directions, and timings provide a compact alternative to application content\[[3](https://arxiv.org/html/2610.00376#bib.bib7),[4](https://arxiv.org/html/2610.00376#bib.bib9)\]\. A classifier trained on these measurements must also confront traffic collected after its training period\. This creates a reason to investigate adaptation through a small labeled context: examples can be supplied to an existing service without fitting new weights\. Such flexibility would be useful if it preserved sufficient accuracy at an acceptable processing time and cost\. This problem is particularly informative for evaluating that proposition\. The observed state is short, the requested answer is one of a finite set of application names, and labeled examples can demonstrate the intended mapping\. These properties make the task straightforward to express through an instruction\-based interface\. Its statistical difficulty remains substantial: anonymous byte counts and time intervals do not explain their relationship to an application name\. Removing domains and other identifiers therefore exposes the question of whether a pretrained service can use numerical traffic evidence, rather than infer a label from an explicit semantic cue\. Suitability for expressing a task and competence at solving it must be evaluated separately\. Jev offers a direct interface for such decisions: a caller supplies a state, a typed question, and candidate answers\[[5](https://arxiv.org/html/2610.00376#bib.bib2)\]\. Its provider emphasizes fast decisions that software can consume\[[6](https://arxiv.org/html/2610.00376#bib.bib3)\]\. General\-purpose GPT APIs also support structured outputs\[[7](https://arxiv.org/html/2610.00376#bib.bib1)\], so output format alone does not distinguish useful classifiers\. The empirical question is whether Jev offers a favorable combination of accuracy, processing time, and cost relative to both a generative service and conventional supervised models\. We study Jev as an observable service, without inferring its internal architecture from its interface\. Our generative comparison uses the GPT\-5\.6 Sol API through Azure\. We compare all main methods on the same numerical prefixes from ten fixed applications in CESNET\-QUICEXT\-25\. Random Forest \(RF\) and Extra Trees \(ET\) are fitted on 8,000 records from collection weeks 1\-4 \(June 1\-28, 2024\), then held fixed alongside Jev’s context for 52,000 records from weeks 5\-30 \(June 29\-December 27, 2024\)\. Extensions on week 5 \(June 29\-July 5, 2024\) examine a larger context and a paired generative\-service comparison\. The application is unknown for each query, but the candidate classes are fixed: the study concerns later\-period classification within a closed set, rather than recognition of unseen classes or a presumed out\-of\-distribution shift\. To the best of our knowledge, this is the first empirical study of general\-purpose decision models, represented here by Jev, for application classification of network flows\. Our longitudinal evaluation uses only the first ten QUIC packets of each flow\. A related Jev IDS project studies intrusion and attack\-category detection from network\-flow features\[[8](https://arxiv.org/html/2610.00376#bib.bib5)\]; our question is which application generated an early QUIC packet prefix\. The study connects three measurements: longitudinal classification quality, improvement from labeled context, and the processing time and API cost of the specified deployments\. Context improves Jev substantially, yet the trained ensembles retain an accuracy advantage in every test week\. On the paired subset, Jev has shorter observed requests and lower API charges than the tested high\-reasoning Sol setting\. Together, these results clarify what adapting a service through examples achieves in this task, and where the remaining accuracy gap limits its practical appeal\. ## 2Background and Related Work Generative and typed\-decision interfaces\.A GPT\-based classifier can receive a task description, measurements, candidate labels, and demonstrations, then return a label in a structured response\. OpenAI’s Structured Outputs supports schema\-constrained responses, including enumerated choices; JSON mode ensures valid JSON but does not itself enforce a particular schema\[[7](https://arxiv.org/html/2610.00376#bib.bib1)\]\. Our Sol configuration uses JSON mode with local label validation\. Jev instead exposes typed questions directly, including a choice over supplied candidates\[[5](https://arxiv.org/html/2610.00376#bib.bib2)\]\. The relevant distinction is how each service exposes and processes a decision, not whether software can consume its output\. We compare the recorded configurations; neither their interfaces nor their names establish architectural speed, correctness, or calibration\. Why evaluate language models on traffic measurements?In\-context learning provides a way to communicate a task through labeled examples without parameter updates\[[1](https://arxiv.org/html/2610.00376#bib.bib14)\]\. Early traffic classification supplies a compact input and a well\-defined output, making it a tractable setting in which to test that mode of adaptation\. However, successful language understanding does not imply a learned correspondence between packet measurements and application identities\. NetLLM addresses the gap between networking inputs and language models through modality encoding, task\-specific output heads, and parameter\-efficient adaptation\[[2](https://arxiv.org/html/2610.00376#bib.bib15)\]\. Its results motivate studying adaptation; they do not establish that serializing numerical prefixes for an unadapted service is sufficient\. TabPFN, a foundation model designed for tabular prediction, further illustrates the importance of the model’s training and representation\[[9](https://arxiv.org/html/2610.00376#bib.bib16)\]\. We evaluate the simpler service\-based route, without extending our findings to these specialized approaches\. Traffic representations and evaluation scope\.CESNET\-QUICEXT\-25 provides longitudinal QUIC measurements\[[3](https://arxiv.org/html/2610.00376#bib.bib7),[10](https://arxiv.org/html/2610.00376#bib.bib8)\], including per\-packet timing, direction, and size\[[4](https://arxiv.org/html/2610.00376#bib.bib9)\]\. Our short numerical prefix makes the observed information comparable across tree ensembles and hosted services\. ET\-BERT instead learns contextual datagram representations through pretraining and task adaptation\[[11](https://arxiv.org/html/2610.00376#bib.bib13)\]\. Its inputs and supervision differ, so its published scores are not numerical baselines for this experiment\. Here ET means Extra Trees\. This study compares methods using the same short numerical prefixes across later collection weeks under the recorded training and service configurations, rather than surveying traffic\-classification methods exhaustively\. Evidence about Jev\.Public documentation describes Jev’s interface and acknowledges uneven capabilities, including numerical tasks\[[5](https://arxiv.org/html/2610.00376#bib.bib2),[12](https://arxiv.org/html/2610.00376#bib.bib4)\]\. A public Jev IDS project uses NSL\-KDD flow features to classify traffic as normal or malicious and to identify attack categories\[[8](https://arxiv.org/html/2610.00376#bib.bib5)\]\. Our task instead identifies one of ten application/service labels from only the first ten QUIC packets, with fixed inputs evaluated across 26 later collection weeks\. A recent survey maps Jev projects across domains rather than benchmarking network application classification\[[13](https://arxiv.org/html/2610.00376#bib.bib6)\]\. Our 52,000\-record comparison measures accuracy, request time, and API charges for the specified configurations; a separate paired subset compares Jev with a generative service\. We keep three properties separate: returning an admissible value, selecting the correct application, and delivering that answer within a useful time and cost\. Only the observed classification and service behavior are tested here; weights, internal computation, and confidence calibration are outside the study\. ## 3Experimental Design ### 3\.1Data and isolation We use a fixed archived revision of a public Parquet conversion of the CESNET DataZoo ORIG distribution\[[14](https://arxiv.org/html/2610.00376#bib.bib10)\]\. The 205 daily files cover nominal June 1\-December 27, 2024 collection partitions and contain 110,511,231 records before filtering\. Sizes and SHA\-256 hashes match the archived mirror\. Additional checks found no differences in packet sequences and timing for 14,750 first\-day records against the official raw source, or in checked IDs, labels, and sequences of 8,000 previously obtained official HDF5 records\. This verifies provenance, not every label’s correctness\. Weeks are seven\-daydaily\-file collection partitions, not asserted UTC flow\-start intervals: a record’s first timestamp can precede its filename date\. Weeks 1\-4 correspond to June 1\-28; week 5 to June 29\-July 5; week 30 to December 21\-27\. Missing dates, October 31 and December 11\-14, are not replaced\. The fixed classes areapple\-privaterelay,facebook\-graph,google\-ads,google\-gstatic,google\-play,google\-services,google\-www,instagram,snapchat, andyoutube\. Labels are inherited from the source; the conversion documents SNI\-derived label provenance\[[14](https://arxiv.org/html/2610.00376#bib.bib10)\]\. Domain strings are not model inputs, although candidate class names are visible as the output vocabulary\. Records must have at least ten packets\. Each input comprises ten triples: inter\-packet time in milliseconds, direction encoded as \-1/\+1, and transport\-payload size in bytes\[[4](https://arxiv.org/html/2610.00376#bib.bib9)\]\. Values are preserved without rescaling\. RF/ET receive a fixed channel\-major vector of 30 float32 values; services receive the same measurements as packet\-major JSON triples with field descriptions\. No test label, domain, address, port, or absolute time is supplied\. A fixed seeded hash of source ID \(seed 2026092305\) prioritizes eligible records, with at most 8,192 candidates per week\. We retain the first 2,000 after excluding previously selected IDs and exact prefix fingerprints\. This operates on raw\-record priorities before deduplication, not uniform sampling over distinct prefix groups, and stops if the pool is insufficient\. Selection is not class\-balanced or prediction\-dependent\. All 60,000 selected training/test IDs and exact float32\-prefix fingerprints are unique\. This does not establish disjoint hosts, users, or underlying connections\. ### 3\.2Models and labeled context RF and ET\.Random Forest\[[15](https://arxiv.org/html/2610.00376#bib.bib11)\]and Extra Trees\[[16](https://arxiv.org/html/2610.00376#bib.bib12)\]each use 500 trees, square\-root feature subsampling, Gini splitting, minimum leaf size one, unlimited depth, no class weighting, and seed 2026092305 in scikit\-learn 1\.7\.2\. They are fitted once on the same 8,000 records and reused without test\-based tuning\. Fitting uses four jobs; timed single\-record prediction uses one job after warm\-up\. These are reference ensembles, not an exhaustive state\-of\-the\-art suite\. A constant reference chooses only the training\-majority class,google\-ads\(1,402/8,000 training records\); its full\-set score is consolidated during manuscript preparation\. Jev\.The requested OpenRouter model istypesafe/jev\-1\.13, checked against returned identifiertypesafe/jev\-1\.13\-20260917\. One choice question requests exactly one application\. Jev\-0 and Jev\-40 denote zero and 40 labeled examples, respectively\. The 40\-example state uses the lowest\-ranked training record in each week/class cell: four examples per class, fixed throughout evaluation, with no weight updates or test feedback\. The 150\-example extension retains the original 40 as its prefix and adds fixed\-ranked training records to reach 15 per class and 37/37/38/38 per training week\. It evaluates the original 2,000 week\-5 records\. One record checks context capacity and remains in the accuracy denominator; the remaining 1,999 use eight workers\. Input counts are 23,614\-23,635 tokens\. These contexts are class\-balanced, unlike the full training set\. GPT\-5\.6 Sol\.The generative baseline is OpenAI GPT\-5\.6 Sol, accessed through Azure\. A fixed salted SHA\-256 ranking of week\-5 sample IDs selects 100 records without consulting labels or predictions\. Deploymentgpt\-5\.6\-sol, with returned identifiergpt\-5\.6\-sol\-2026\-07\-09, useshighreasoning effort, a 4,096\-completion\-token cap, non\-streaming JSON output, and a 60 s request deadline\. Sol\-0 and Sol\-40 denote this same configuration with zero or the original 40 examples\. Inputs, candidates, and examples match Jev; the API/output wrapper differs\. Serial requests alternate arm order by sample\. Both services and RF/ET are compared on this identical subset\. It is exploratory: week 5 had already been examined, and this is neither a blind new holdout nor an optimized LLM configuration search\. ### 3\.3Metrics and operational accounting Accuracy uses the fixed planned denominator\. Macro\-F1 averages the ten per\-class F1 scores, assigning zero to undefined scores\. Final coverage is complete: 52,000 per main arm, 2,000 for Jev\-150, and 100 per Sol arm\. The smaller studies overlap the main test set\. Local time encloses one warm\-startedpredictcall\. Remote time includes client/worker communication, network, gateway, and service processing for a successful request\. Neither includes observing the ten packets\. Later\-week ML batch times are not presented as single\-flow latency; paired timing uses the original week\-5 single\-record measurements\. Jev uses eight workers per main arm with partially overlapping execution; Sol is serial at different times over a different service path\. Medians and 95th percentiles therefore compare deployments, not isolated architectures under matched load\. Interrupted runs preserve successful predictions and resume only missing queries with unchanged settings\. Each query stops at its first valid answer, irrespective of correctness\. Failures, retries, and unknown\-charge attempts remain in accounting\. Successful\-attempt time excludes retry waits and does not measure first\-attempt availability\. Jev cost comes from returned usage\. Sol uses conservative bookkeeping rates of USD 6\.25/M input and USD 30/M output tokens without cache discounts, not a verified invoice or current\-tariff claim\. Conservative totals additionally retain a 25% allowance and failed/unknown\-request reservations\. Local RF/ET have no API fee; hardware, energy, and labor are unpriced\. The September 23, 2026 runs used Windows 11, Python 3\.12\.14, NumPy 2\.3\.5, and an Intel family\-6/model\-158 CPU with 12 logical processors, without a local GPU\. Frozen inputs, model identifiers, response hashes, and labels were locally cross\-checked; aggregate data accompany this paper\. ## 4Results ### 4\.1Full\-set accuracy and prediction behavior [Table 1](https://arxiv.org/html/2610.00376#S4.T1)reports the same 52,000 records\. Forty examples improve Jev by 18\.62 percentage points, but leave a 41\.53\-point gap to RF\. Macro\-F1 preserves the ordering among RF, ET, and the two Jev configurations\. The training\-majority constant reaches 22\.79% accuracy: Jev\-40 exceeds it by 5\.62 points and has substantially higher macro\-F1, whereas zero\-shot Jev is below it on both metrics\. Table 1:Full test set, with 100% final valid coverage for every method\. Accuracy is percent; supervision budgets differ\.Zero\-shot predictions concentrate ongoogle\-www\(45,010\),youtube\(6,958\), andgoogle\-gstatic\(32\), with no other predicted class\. This is not uniform random guessing across balanced classes\. Context helps unevenly: Jev\-40 recalls 98\.3% of 815apple\-privaterelayrecords, close to RF’s 99\.4%, but only 15\.8% ofgoogle\-adsand 16\.9% ofgoogle\-play, versus RF’s 78\.3% and 74\.2%\. These differences diagnose the aggregate result; favorable classes are not selected to replace it\. ### 4\.2Across collection weeks RF and ET remain above Jev\-40 in all 26 test weeks \([Figure 1](https://arxiv.org/html/2610.00376#S4.F1)\)\. RF changes from 72\.70% in week 5 to 69\.50% in week 30, and ET from 70\.00% to 64\.85%\. Both reach their minima in week 24 \(63\.60% and 58\.45%\) and then recover\. Jev\-40 peaks at 37\.60% in that week without overtaking either ensemble\. There is no observed monotonic collapse or accuracy crossover\. Changing class proportions, traffic behavior, and missing dates may contribute; chronological separation alone does not establish an out\-of\-distribution advantage\. Figure 1:Weekly accuracy on 2,000 identical test records per week, with fixed models and context\. These are descriptive collection\-week trajectories\. ### 4\.3More labeled examples On the same 2,000 week\-5 records, Jev with zero, 40, and 150 examples correctly classifies 206, 518, and 690 records, respectively: 10\.30%, 25\.90%, and 34\.50% \([Figure 2](https://arxiv.org/html/2610.00376#S4.F2)\)\. The 150\-example arm has macro\-F1 0\.4560 and remains 38\.20 accuracy points below RF\. Gains of 15\.60 and 8\.60 points show that this fixed context family helps, but three points from one week do not establish a scaling law\. All 2,000 requests succeed; the 1,999\-record eight\-worker stage has a 0\.910 s median\. Its single\-worker capacity probe is excluded only from that timing summary\. No matched\-40/150\-example tree experiment has yet been run\. Figure 2:Week\-5 accuracy versus Jev context size on the same 2,000 records\. The RF/ET reference lines use 8,000 training records, not equal label budgets\. ### 4\.4Paired generative\-service comparison [Table 2](https://arxiv.org/html/2610.00376#S4.T2)uses the same 100 week\-5 records throughout\. Sol\-40 correctly classifies 37 versus Jev\-40’s 29\. Sol alone is correct on 19 records, Jev alone on 11, and both on 18\. An exploratory two\-sided exact McNemar test gives p = 0\.2005: this small sample does not establish stable superiority or equivalence, and shared collection conditions limit population\-level interpretation\. RF/ET achieve 65/70; their reversed ordering relative to the full set also cautions against replacing the main experiment with this subset\. Table 2:Paired week\-5 subset \(100 records\)\. Suffixes denote context examples\. Sol is OpenAI GPT\-5\.6 Sol through Azure, with high reasoning effort\. P50/P95 are successful prediction/request times in seconds\. RF/ET are local; Jev/Sol have different concurrency and service paths\.The 40\-example request medians differ by 5\.29 s, a ratio of about 8\.0 \([Figure 3](https://arxiv.org/html/2610.00376#S4.F3)\)\. This describes service configurations, not architectural speedup\. Full\-set Jev\-0/40 medians are 0\.716/0\.773 s, with P95 of 0\.832/0\.911 s\. Sol’s usage records help interpret its longer response time: zero\-shot median completion/reasoning counts are 1,049/1,024 tokens, versus 368\.5/350\.5 with context, despite prompt medians increasing from 260 to 5,676 tokens\. Longer context coincides with less reasoning and shorter time, but does not causally isolate this mechanism\. Lower reasoning effort was not tested\. Figure 3:Observed accuracy and median successful processing time on the paired 100 records\. Squares are local ML, circles Jev, and triangles OpenAI GPT\-5\.6 Sol with high reasoning effort through Azure\. Suffixes \-0/\-40 denote context examples\. Timing boundaries differ\. ### 4\.5Failures and cost Sol obtains 200 valid labels over 211 attempts\. Eleven failures comprise two client permission errors, seven deadlines, one invalid length\-limited response, and one HTTP error\. One query needs four attempts; its first valid response is incorrect and retained\. Successful\-request durations sum to 54\.6 minutes, or 62\.9 including failures, excluding inter\-run gaps\. These sums are not total experiment wall\-clock time\. Jev\-0/40 record 52,008/52,020 started attempts, retaining terminal failures and 7/15 starts without terminal records after interruption\. Unknown charges are not treated as zero\. Table 3:Batch cost in USD\. Jev known cost is API\-reported; conservative totals include margins and failed/unknown reservations\. Sol uses budget\-rate estimates, not an invoice\.The conservative batch total is USD 36\.4150 \([Table 3](https://arxiv.org/html/2610.00376#S4.T3)\), excluding earlier pilots and research\. For 100 paired successful responses, Jev\-0/40 report USD 0\.002341/0\.028234; Sol estimates are USD 3\.927254/5\.613564 before failures and allowances\. These are service\-price observations, not invoice\-verified or energy\-efficiency comparisons\. ## 5Discussion and Limitations What the accuracy gap means\.The task is easy to state but requires a mapping that its numerical inputs do not make explicit\. Concentrated zero\-shot predictions and gains from labeled context suggest that examples help supply this missing task information, although they do not reveal how either service reasons internally\. Increasing Jev’s context improves accuracy without closing the gap to the trained ensembles\. The remaining gap may reflect the representation, example count and selection, service capability, or several of these together; this design cannot separate them\. It does not imply that language models are unsuitable for language\-rich logs, network intents, or other representations\. RF/ET use 8,000 training records, whereas Jev receives 40/150 class\-balanced examples and Sol at most 40; these results compare practical configurations, not intrinsic label efficiency\. No matched\-example tree control was run, so this study cannot estimate intrinsic label efficiency\. What the service comparison means\.Jev’s shorter requests and lower charges make its decision interface worth examining, but operational efficiency is useful only at an acceptable error rate\. The paired subset does not resolve an accuracy advantage for either hosted service\. Moreover, high\-reasoning Sol is one configuration rather than an optimized speed baseline\. API wrappers, concurrency, execution times, network paths, caching, and recovery differ, while local ML includes Python prediction overhead\. We do not measure packet\-observation time, feature extraction, live deadline satisfaction, or local Jev inference\. The measured trade\-off therefore concerns these deployments; it establishes neither real\-time suitability nor an architectural speed advantage\. Limits of the evidence\.One network, ten classes, one weekly selection, and one fixed context family limit generalization\. Earlier exploratory work influenced scope, and week\-5 extensions reuse examined data\. Exact ID/prefix isolation does not exclude near duplicates or shared host/session conditions; hosted\-model pretraining overlap is unknown\. Labels are inherited from the dataset\. Weekly trajectories describe later\-period performance, but correlated flows and chronological separation alone cannot identify a causal drift mechanism or establish behavior in independent environments\. Reproducibility and next steps\.The results apply to the recorded model identifiers and execution date; services and prices may change\. Aggregate data and frozen\-source hashes support the tables and figures, without constituting a public release of every raw response\. Multiple example selections, independent networks, and lower\-reasoning Sol settings could test how widely these observations hold\. These extensions have not been run\. Until then, the evidence supports a bounded conclusion: examples help adapt this decision service, while the accuracy needed to replace the tested supervised classifiers remains unachieved\. ## 6Conclusion and Future Work Early traffic classification provides a concrete test of turning a pretrained service into a classifier through instructions and examples\. On the ten\-application numerical\-prefix task studied here, labeled context improves Jev, including a further gain from 40 to 150 examples on week 5, but RF/ET trained on 8,000 records remain more accurate in every later collection week\. Jev’s shorter requests and lower charges relative to one high\-reasoning Sol configuration reveal an operational trade\-off, with the paired accuracy difference unresolved\. In this setting, a convenient decision interface did not yield accuracy comparable to the trained ensembles\. Future work should test whether Jev can make better use of early\-packet evidence through alternative representations and systematically selected examples\. Matched\-label\-budget tree controls would clarify how much of the observed gap reflects supervision, while evaluation on independent networks and collection periods would test generality\. Comparing service configurations under aligned timing and cost conditions would further establish where decision services may be useful for traffic classification\. These directions explore Jev’s potential without assuming an advantage that the present results do not show\. ## References - \[1\]T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan,et al\.\(2020\)Language models are few\-shot learners\.Note:arXiv:2005\.14165External Links:[Link](https://arxiv.org/abs/2005.14165)Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p1.1),[§2](https://arxiv.org/html/2610.00376#S2.p2.1)\. - \[2\]D\. Wu, X\. Wang, Y\. Qiao, Z\. Wang, J\. Jiang, S\. Cui, and F\. Wang\(2024\)NetLLM: adapting large language models for networking\.InProceedings of the ACM SIGCOMM 2024 Conference,pp\. 661–678\.External Links:[Document](https://dx.doi.org/10.1145/3651890.3672268)Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p1.1),[§2](https://arxiv.org/html/2610.00376#S2.p2.1)\. - \[3\]J\. Luxemburk, K\. Hynek, and J\. Mücke\(2025\)CESNET\-QUICEXT\-25: a year\-long QUIC traffic dataset with extended attributes\.Note:Zenodo, version 1External Links:[Document](https://dx.doi.org/10.5281/zenodo.17249078),[Link](https://zenodo.org/records/17249078)Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p2.1),[§2](https://arxiv.org/html/2610.00376#S2.p3.1)\. - \[4\]CESNET\(2026\)CESNET DataZoo: data features\.Note:[https://cesnet\.github\.io/cesnet\-datazoo/features/](https://cesnet.github.io/cesnet-datazoo/features/)Accessed September 24, 2026Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p2.1),[§2](https://arxiv.org/html/2610.00376#S2.p3.1),[§3\.1](https://arxiv.org/html/2610.00376#S3.SS1.p4.1)\. - \[5\]TypeSafe AI\(2026\)Introduction: jev and typed decisions\.Note:[https://docs\.typesafe\.ai/introduction](https://docs.typesafe.ai/introduction)Accessed September 24, 2026Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p4.1),[§2](https://arxiv.org/html/2610.00376#S2.p1.1),[§2](https://arxiv.org/html/2610.00376#S2.p4.1)\. - \[6\]D\. Almeida\(2026\)Introducing system one models & jev\.Note:TypeSafe AI,[https://typesafe\.ai/blog/introducing\-system\-one\-models\-and\-jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)September 15, 2026; provider descriptionCited by:[§1](https://arxiv.org/html/2610.00376#S1.p4.1)\. - \[7\]OpenAI\(2026\)Structured model outputs\.Note:[https://developers\.openai\.com/api/docs/guides/structured\-outputs](https://developers.openai.com/api/docs/guides/structured-outputs)Accessed September 28, 2026Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p4.1),[§2](https://arxiv.org/html/2610.00376#S2.p1.1)\. - \[8\]Jev IDS contributors\(2026\)Jev IDS: intrusion detection with a system one model\.Note:GitHub repository,[https://github\.com/jev\-ids/jev\-ids](https://github.com/jev-ids/jev-ids)Pinned commit 57fa123886a6d08b223b03e340395ee42789d08a; accessed September 28, 2026Cited by:[§1](https://arxiv.org/html/2610.00376#S1.p6.1),[§2](https://arxiv.org/html/2610.00376#S2.p4.1)\. - \[9\]N\. Hollmann, S\. Müller, L\. Purucker, A\. Krishnakumar, M\. Körfer, S\. B\. Hoo, R\. T\. Schirrmeister, and F\. Hutter\(2025\)Accurate predictions on small data with a tabular foundation model\.Nature637,pp\. 319–326\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08328-6)Cited by:[§2](https://arxiv.org/html/2610.00376#S2.p2.1)\. - \[10\]J\. Mücke, M\. Nawrocki, R\. Hiesgen, P\. Sattler, J\. Zirngibl, G\. Carle, J\. Luxemburk, T\. C\. Schmidt, and M\. Wählisch\(2025\)Waiting for QUIC: passive measurements to understand QUIC deployments\.Proceedings of the ACM on Networking3\(CoNEXT4\),pp\. 41:1–41:26\.External Links:[Document](https://dx.doi.org/10.1145/3768988)Cited by:[§2](https://arxiv.org/html/2610.00376#S2.p3.1)\. - \[11\]X\. Lin, G\. Xiong, G\. Gou, Z\. Li, J\. Shi, and J\. Yu\(2022\)ET\-BERT: a contextualized datagram representation with pre\-training transformers for encrypted traffic classification\.Note:arXiv:2202\.06335External Links:[Link](https://arxiv.org/abs/2202.06335)Cited by:[§2](https://arxiv.org/html/2610.00376#S2.p3.1)\. - \[12\]TypeSafe AI\(2026\)Jev 1\.13 jaggedness\.Note:[https://docs\.typesafe\.ai/model\-jaggedness/jev\-1\.13](https://docs.typesafe.ai/model-jaggedness/jev-1.13)Accessed September 24, 2026Cited by:[§2](https://arxiv.org/html/2610.00376#S2.p4.1)\. - \[13\]G\. Ling, M\. Xue, and Z\. Ye\(2026\)Jev in the Wild: a data\-driven analysis of the Jev model’s functionality, applications and ecosystem\.Note:arXiv:2609\.30216External Links:[Link](https://arxiv.org/abs/2609.30216)Cited by:[§2](https://arxiv.org/html/2610.00376#S2.p4.1)\. - \[14\]Lystea\(2026\)CESNET\-QUICEXT\-25: canonical flow parquet\.Note:Hugging Face dataset,[https://huggingface\.co/datasets/Lystea/CESNET\-QUICEXT25\-PARQUET](https://huggingface.co/datasets/Lystea/CESNET-QUICEXT25-PARQUET)Archived metadata and field\-level source checks retainedCited by:[§3\.1](https://arxiv.org/html/2610.00376#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2610.00376#S3.SS1.p3.1)\. - \[15\]L\. Breiman\(2001\)Random forests\.Machine Learning45,pp\. 5–32\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1010933404324)Cited by:[§3\.2](https://arxiv.org/html/2610.00376#S3.SS2.p1.1)\. - \[16\]P\. Geurts, D\. Ernst, and L\. Wehenkel\(2006\)Extremely randomized trees\.Machine Learning63,pp\. 3–42\.External Links:[Document](https://dx.doi.org/10.1007/s10994-006-6226-1)Cited by:[§3\.2](https://arxiv.org/html/2610.00376#S3.SS2.p1.1)\.
相似文章
用于文本分类的语言模型:从词袋模型到Jev AI
本文概述了文本分类方法的历史演进,从词袋模型到近期发布的Jev AI模型,该模型提供了高效且通用的分类能力。
@NathanFlurry:无炒作的JEV解释:JEV并不取代GPT/Claude,JEV只是一个*非常*智能的switch语句,比如如果2…
Diogo Almeida发布了JEV,这是一个新的AI模型,它作为智能switch语句处理分类和路由等任务,声称在速度和效率上显著优于现有模型。
Jev 的校准有多准确?
一篇独立博客对 TypeSafe 的 Jev "System One" 分类模型的分析,通过一些已被充分理解的问题,测试其预测的概率分布与已知真实分布(例如麦克斯韦-玻尔兹曼速度分布)的匹配程度。文章结论是,尽管评估成本极低,Jev 的校准仍然很差。
@svpino: Jev 太厉害了!如果你还不知道,Jev 是一个新的 "System One" 模型,针对决策进行了优化。例如…
Jev 是一个新的 AI 模型,针对分类任务中的快速且低成本决策进行了优化,通过一个使用 Apify 和 GPT-5-mini 的酒店评论分类器项目进行了演示。
Jev 不是 LLM 杀手,也不仅仅是一个分类器。我们将其部署于生产环境并与真实用户一起使用。以下是我们学到的经验。
文章分享了使用 Jev(一个语义决策引擎)的生产见解,它通过与 LLMs 并行高效处理常规决策来增强 AI 代理系统,而不替代生成模型。