ornith-ai/Ornith-1.5-9B

Hugging Face Models Trending 模型

摘要

Ornith-1.5-9B 是一款9B参数密集型AI模型,通过端到端自我改进来推进基础模型构建,并利用强化学习优化任务生成、框架构建和解决方案部署。它在与Qwen和Gemma等其他模型相比时,在各种基准测试上表现出竞争力性能。

任务:文本生成 标签:transformers、safetensors、qwen3_5、image-text-to-text(图文转文本)、text-generation(文本生成)、conversational(对话式)、license:mit(MIT许可证)、eval-results(评估结果)、endpoints_compatible(端点兼容)、region:us(区域:美国)
查看原文
查看缓存全文

缓存时间: 2026/08/23 03:59

ornith-ai/Ornith-1.5-9B · Hugging Face

来源:https://huggingface.co/ornith-ai/Ornith-1.5-9B 叽喳叽喳!🐦 我们推出 Ornith-1.5,这是通过端到端自我改进构建基础模型的重要一步。Ornith-1.5 扩展了 Ornith-1.0(该模型在 Qwen3.5 和 Gemma4 基础上通过额外的持续预训练、中期训练和后期训练开发而来),将自我改进循环从脚手架和展开优化扩展到联合优化任务生成、脚手架构建和解决方案展开。Ornith-1.5 不依赖于固定的人为策划任务集和手动设计的框架,而是持续生成新的训练任务,发现解决它们的有效策略,并通过强化学习改进策略。有关任务、框架和展开奖励设计的更多细节,请参考我们的博客 (https://ornith.ai/ornith_1_5.html)。

Ornith 1.5 9B 基准测试结果

Ornith 1.5 9B

本文档记录了 Ornith-1.5-9B,Ornith-1.5 系列中最轻量级的成员——一个 9B 的稠密模型,旨在高效进行单 GPU 部署,并可通过其量化版本 Ornith-1.5-9B-Mobile 在移动设备上部署。

基准测试

基准测试Ornith-1.5-9BOrnith-1.0-9BQwen3.5-9BQwen3.6-35B-A3BGemma-4-31B
编码
Terminal-Bench 2.1 (Terminus-2)46.243.121.352.542.1
Terminal-Bench 2.1 (Claude Code)4740.618.949.2-
SWE-bench Verified70.669.453.273.452
SWE-bench Pro47.542.931.349.535.7
SWE-bench Multilingual54.45239.767.251.7
NL2Repo32.427.216.229.415.5
SWE Atlas - QnA20.617.99.215.5-
推理
HLE (no tools)20.216.814.721.419.5
HLE (with tools)30.526.424.528.926.5
GPQA Diamond86.482.581.78684.3
智能体
MCP-Atlas54.249.446.862.855
Toolathlon-Verified41.233.429.641.752.8
WideSearch59.555.853.660.154.2
BrowseComp56.444.841.562-
ClawEval66.563.153.268.748.5
  • 所有为 Ornith-1.5 报告的结果均为五次独立运行的平均值。
  • Terminal-Bench 2.1 (Terminus-2): 我们使用 Harbor/Terminus-2 框架评估 Terminal-Bench 2.1,设置为 parser=json,temperature=1.0,top_p=1.0,上下文窗口为 128K。每次运行使用 4 小时超时,32 个 CPU 核心和 48GB 内存,结果为 5 次运行的平均值。我们调整了 Qwen 聊天模板以确保训练和推理之间的一致性 (https://huggingface.co/ornith-ai/Ornith-1.5-9B/blob/main/chat_template.jinja),并修改了 Harbor 以对齐 vLLM 的 reasoning_content 键。
  • Terminal-Bench 2.1 (Claude Code): 我们使用 Claude Code 2.1.126 评估 Terminal-Bench 2.1,设置为 parser=json,temperature=1.0,top_p=1.0,max_new_tokens=131072。结果为 5 次运行的平均值。同样,Qwen 聊天模板需要修改。
  • SWE-Bench Verified, Pro and Multilingual: 使用 OpenHands 框架,设置 temp=1.0,top_p=0.95,上下文窗口 256k。评估全程应用防黑客保护措施:从本地仓库镜像中移除 Git 历史记录以防止访问先前的解决方案或提交;禁用网络访问,防止模型检索外部信息或资源。
  • DeepSWE: 使用 Claude Code 框架评估,设置 temperature=1.0,top_p=0.95,上下文窗口 256K。
  • SWE Atlas QnA: 使用 mini SWE agent 框架,设置 temp=1.0,top_p=0.95,上下文窗口 128K。结果为 5 次运行的平均值。
  • NL2Repo: 设置 temperature=1.0,top_p=1.0,上下文 400K,输出 48K。为防止奖励黑客,屏蔽了对指定 GitHub 仓库和 pip 包的访问。
  • HLE: 使用 Claude 4.6 Opus 作为评判模型进行评估。
  • MCP-Atlas: 所有模型均在思考模式下于 500 个任务的公共子集上进行评估,每个任务超时 10 分钟。我们使用 Claude 4.8 Opus 作为评判模型。
  • Toolathlon-Verified: 我们使用官方评估服务,最大 token 限制设置为 128K。
  • ClawEval: 一个基于真实用户任务分布的智能体代码基准测试;temp=0.6,上下文 256K。

快速入门

📝 注意: Ornith-1.5-9B 是一个推理模型:默认情况下,助手回复会在最终答案前以一个 ...</think> 块开始。下面的部署方案启用了推理解析器,以便将思维链返回到单独的 reasoning_content 字段中,并启用了工具调用解析器,以便将模型的 `` 块作为 OpenAI 风格的 tool_calls 呈现。

部署 Ornith-1.5-9B 需要较新的运行时:

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.19.1
  • SGLang ≥ 0.5.9

推荐采样参数:

  • 通用任务: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • 精确编码任务: temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

部署 Ornith-1.5-9B

Ornith-1.5-9B 是一个约 9B 的稠密模型(bf16 格式约 19 GB),因此可以在单个 80GB GPU 上部署。下面的方案启动一个兼容 OpenAI 的服务器;如果你想跨多个 GPU 分片,可以添加 --tensor-parallel-size/--tp

  • vLLM

    vllm serve ornith-ai/Ornith-1.5-9B \
      --served-model-name Ornith-1.5-9B \
      --host 0.0.0.0 --port 8000 \
      --max-model-len 262144 \
      --gpu-memory-utilization 0.90 \
      --enable-prefix-caching \
      --enable-auto-tool-choice --tool-call-parser qwen3_xml \
      --reasoning-parser qwen3 \
      --trust-remote-code
    
  • SGLang

    python -m sglang.launch_server \
      --model-path ornith-ai/Ornith-1.5-9B \
      --served-model-name Ornith-1.5-9B \
      --host 0.0.0.0 --port 8000 \
      --context-length 262144 \
      --mem-fraction-static 0.85 \
      --tool-call-parser qwen3_coder \
      --reasoning-parser qwen3
    
用于长上下文

Ornith-1.5-9B 支持最大 262,144 个 token 的上下文窗口。当任务的组合输入和输出必须超出此限制时,我们建议使用 RoPE 缩放来扩展有效窗口——YaRN 是我们验证过的技术,并且已内置于 vLLM 和 SGLang 中。缩放因子为 4.0 时,可用窗口扩展到大约 100 万 token。

你可以通过以下两种方式之一启用 YaRN:

  • 编辑检查点的 config.json 在模型配置中添加 rope_scaling 块:

    {
      "rope_scaling": {
        "rope_type": "yarn",
        "factor": 4.0,
        "original_max_position_embeddings": 262144
      }
    }
    
  • 在启动时覆盖。 保持检查点不变,在上面的部署命令中添加等效的标志。

    vLLM:

    VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ornith-ai/Ornith-1.5-9B ... \
      --hf-overrides '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
      --max-model-len 1000000
    

    SGLang:

    SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... \
      --json-model-override-args '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
      --context-length 1000000
    

📝 注意: 开源运行时静态实现 YaRN:无论请求长度如何,都应用相同的缩放因子,这可能会略微损害普通长度输入的质量。仅在你的工作负载确实需要更长窗口时才启用 rope_scaling,并将 factor 设置为匹配它——目标窗口大约是 factor × 262,144,因此如果你的请求上限在 524,288 个 token 左右,factor: 2.0 是更好的设置。

通过 Chat Completions API 使用 Ornith-1.5-9B

一旦 vLLM 或 SGLang 服务器运行起来,你可以使用任何兼容 OpenAI 的客户端与之对话。

基本用法
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",  # 任何非空字符串对本地服务器都有效
)

response = client.chat.completions.create(
    model="Ornith-1.5-9B",
    messages=[
        {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
    ],
    temperature=0.6,
    top_p=0.95,
    max_tokens=1024,
)

message = response.choices[0].message
# reasoning_content 保存推理过程;content 保存最终答案。
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

你还可以流式传输 token,或者为模型提供工具——Ornith-1.5-9B 会发出格式正确的函数调用,服务器将其解析为标准的 tool_calls 字段:

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

response = client.chat.completions.create(
    model="Ornith-1.5-9B",
    messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.6,
    max_tokens=2048,
)

tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
# -> get_weather {"city": "Paris"}

你可以将任何兼容 OpenAI 的 SDK(Python、Node.js 等)或 curl 指向同一个 /v1/chat/completions 端点。

智能体用法

Ornith-1.5-9B 暴露了一个兼容 OpenAI 的工具调用端点,它开箱即用地支持标准智能体框架。

使用 Ornith 与智能体的示例:

Ollama

ollama run ornith-1.5:9b

Atomic.chat

# 两个运行时都加载 Ornith 的 GGUF 构建版本(在 ornith-ai/Ornith-1.5-9B-GGUF 发布)。
# llama.cpp — 在端口 8000 上提供兼容 OpenAI 的 API。
llama-server -hf hf.co/ornith-ai/Ornith-1.5-9B-GGUF --port 8000 -c 262144

llama.cpp

# 两个运行时都加载 Ornith 的 GGUF 构建版本(在 ornith-ai/Ornith-1.5-9B-GGUF 发布)。
# llama.cpp — 在端口 8000 上提供兼容 OpenAI 的 API。
llama-server -hf hf.co/ornith-ai/Ornith-1.5-9B-GGUF --port 8000 -c 262144

Hermes Agent

# Hermes 可与任何兼容 OpenAI 的端点对话——将其指向你的 Ornith 服务器。
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="ornith-ai/Ornith-1.5-9B"

OpenClaw

# OpenClaw 可与任何兼容 OpenAI 的端点对话——将其指向你的 Ornith 服务器。
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="ornith-ai/Ornith-1.5-9B"

Unsloth Studio

pip install unsloth

# 加载 Ornith 用于快速本地推理或微调(Python):
# from unsloth import FastLanguageModel
# model, tokenizer = FastLanguageModel.from_pretrained(
#     "unsloth/Ornith-1.5-9B-GGUF",
#     max_seq_length=262144,
#     load_in_4bit=True,
# )

编码 CLI

Ornith-1.5-9B 针对基于终端的编码智能体进行了优化。将任何兼容 OpenAI 的编码 CLI 指向你的 Ornith-1.5-9B 端点(设置 OPENAI_BASE_URLOPENAI_API_KEY),以理解大型代码库、自动化繁琐工作并更快地交付。

OpenCode
# 将你的本地 Ornith 端点注册为 ~/.config/opencode/opencode.json 中的提供者:
# {
#   "$schema": "https://opencode.ai/config.json",
#   "provider": {
#     "ornith": {
#       "npm": "@ai-sdk/openai-compatible",
#       "name": "Ornith (local)",
#       "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
#       "models": { "ornith-ai/Ornith-1.5-9B": { "name": "Ornith-1.5-9B" } }
#     }
#   }
# }

opencode

引用

如果你觉得我们的工作有帮助,请随时引用我们。

@misc{ornith_1_5,
  title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
  url = {https://ornith.ai/ornith_1_5.html},
  author = {{Ornith Team}},
  year = {2026}
}

相似文章

ornith-ai/Ornith-1.5-9B-GGUF

Hugging Face Models Trending

Ornith-1.5是一款新的AI模型,它通过强化学习引入了端到端的自我改进,针对单GPU和边缘设备的高效部署进行了优化。

ornith-ai/Ornith-1.5-35B-A3B

Hugging Face Models Trending

Ornith-1.5-35B-A3B 是一种混合专家AI模型,每个令牌仅激活30亿参数,并在编码和代理基准测试中超越了类似规模的模型,如Qwen和Gemma。

ornith-ai/Ornith-1.5-35B-A3B-GGUF

Hugging Face Models Trending

Ornith-1.5-35B-A3B 是一款新的AI基础模型,通过采用端到端自我改进,在编码和代理基准测试上取得了卓越的性能,每个令牌仅激活约30亿个参数。

deepreinforce-ai/Ornith-1.0-9B

Hugging Face Models Trending

deepreinforce-ai 发布了 Ornith-1.0,一个开源编码代理模型系列,在编码基准测试上实现了最先进的性能,提供从 9B 到 397B 的参数规模,采用自我改进训练框架和 MIT 许可证。