ggdxwz
V2EX  ›  OpenAI

下一代 GPT5.6 为了在跑分上作弊,自主挖掘零日漏洞从沙盒逃逸,然后把 Hugging Face 黑了

  •  
  •   ggdxwz · 4h 8m ago · 4372 views

    这几天的💩终于串起来了

    简单来讲就是 OAI 一个内部模型 (GPT 6 / 5.6 Sol+) 为了能在 ExploitGym 测试上获得更高分数,找到软件包的缓存代理中的零日漏洞,自己提权从沙盒里面跑出来,上网发现抱抱脸可能有这个测试的数据集,然后把 Hugging Face 的生产服务器黑了,最终拿到了这个测试的答案🤔

    最搞笑的是 Hugging Face 用 GPT 5.6 来防结果没有 Cyber 权限被拒,最后只能用自己部署的 GLM 5.2 才勉强解决问题

    现在 Trump 说太危险了,能救场的中国模型都得上 Ban 位😅

    这是 OpenAI 的 PR 稿: https://openai.com/index/hugging-face-model-evaluation-security-incident/

    这是 Hugging Face 的报告: https://huggingface.co/blog/security-incident-july-2026

    全文翻译图片在这里,长图就不整个贴出来了:

    https://i.imgur.com/8pGTKn6.jpeg

    35 replies    2026-07-22 11:52:20 +08:00
    MIUIOS
        1
    MIUIOS  
       4h 4m ago
    claude 和 gpt 都是存在直接抄答案的嫌疑,现在的跑分真的看看就好了,还得实际体验下模型才知道
    zx9481
        2
    zx9481  
       3h 43m ago   ❤️ 1
    感觉是自导自演的炒作
    cat9life
        3
    cat9life  
       3h 33m ago
    不太懂,是 oai 在评估,还是抱脸
    Rehtt
        4
    Rehtt  
       3h 29m ago
    @cat9life oai 在评估新模型跑分,结果新模型把 hf 黑了拿到题目答案
    ggdxwz
        5
    ggdxwz  
    OP
       3h 25m ago   ❤️ 1
    @zx9481 #2 不太像,抱抱脸没理由演戏,感觉他们实打实出问题了,我的 space 那段时间莫名其妙炸了
    OAI 要是敢为了刷分去黑数据集,今年 IPO 可以取消把钱拱手让给 A\
    wonderfulcxm
        6
    wonderfulcxm  
       3h 19m ago via iPhone
    说明 llm 可以为了达到目的不择手段,那挺可怕的。
    xingzhi95
        7
    xingzhi95  
       3h 15m ago
    现在顶级 AI 在安全领域有神一般的能力,根本就不是人类可以抗衡的
    YanSeven
        8
    YanSeven  
       3h 15m ago via Android
    这玩意离谱到稍微改编一下,就是电影情节。
    wonderfulcxm
        9
    wonderfulcxm  
       3h 7m ago via iPhone
    glm5.2 并不是防御,只是用来分析日志和取证的
    Rorysky
        10
    Rorysky  
       3h 6m ago
    听上去就是广告
    HeyWeGo
        11
    HeyWeGo  
       3h 5m ago
    这是不是类似竞赛书里说的那个只验证结果的,直接 print 结果就行那种思路?
    dingawm
        12
    dingawm  
       2h 56m ago
    @Rorysky #10 这个对于 HF 也许是广告,但是对于 OpenAI 感觉更是个负面事件啊,不更强化了开源模型的叙事吗?如果这个事件里的防御方强调是 GPT ,那确实更像是个广告
    dabbit
        13
    dabbit  
       2h 41m ago
    蛮吓人的,有点不受控了。
    huanxianghao
        14
    huanxianghao  
       2h 41m ago
    看着我项目里的这个问题,gpt 5.6 怎么都解决不了,我陷入了沉思
    ggdxwz
        15
    ggdxwz  
    OP
       2h 37m ago
    @huanxianghao #14 这个里面用的都是调整过的,猜测其中 5.6 应该只是做 subagent ,主线程是没有名字的下一代旗舰
    noahliaszn
        16
    noahliaszn  
       2h 12m ago
    模型为了答案都不择手段 有点恐怖了
    iosyyy
        17
    iosyyy  
       1h 59m ago
    你这前面和后面都不搭边, Hugging Face 的意思明显是商用模型(没有指向 5.6)有安全限制. 所以他们最开始用的就是开源权重的 GLM 5.2
    iosyyy
        18
    iosyyy  
       1h 57m ago
    @iosyyy 准确的说是评估问题的时候出现的, 感觉这个明显是吹 openai 模型牛逼的广告. 全文只是后面提了一嘴开源模型
    wsseo
        19
    wsseo  
       1h 45m ago
    天天炒作。
    ggdxwz
        20
    ggdxwz  
    OP
       1h 37m ago
    @iosyyy #17 你要是前几天有关注就知道,抱抱脸两家的旗舰都试了,每一个能过 cyber ,然后文章里面也写了,但没点名,你可以自己核对
    https://huggingface.co/blog/security-incident-july-2026#the-asymmetry-problem
    dbskcnc
        21
    dbskcnc  
       1h 25m ago
    ai 真的挺恐怖的,现在能力才这样稍不慎就能搞事,要是以后能力升级,然后真成了脱缰的魔兽时候,人挡杀人,佛挡杀佛,特别是 it 越来越发达,它在数字世界可真像现在的游戏一样,随手就干,大家都得升级装备才行了
    shintendo
        22
    shintendo  
       1h 22m ago
    @iosyyy “所以他们最开始用的就是开源权重的 GLM” 你读一下报告呢

    fate
        23
    fate  
       1h 20m ago
    感觉扯犊子的可能性更大,如果这事情是真的,那么 openai 内部的安全和入侵检测做的就像是屎一样,毕竟模型可是在他们内网横向并找到了一个出网的机器。
    lichwong
        24
    lichwong  
       1h 18m ago
    怕以后进化到为了解决某些问题黑进军方直接发射核弹
    catazshadow
        25
    catazshadow  
       1h 17m ago via Android
    碟中谍原来是纪录片
    litmxs
        26
    litmxs  
       1h 15m ago via iPhone
    想起之前看到的,让大模型执行需要 root 的任务但是没有给 root 权限,大模型自己想办法提权了(当前账号没有 root 但是在 docker 用户组里面,通过创建容器提权限)
    xing7673
        27
    xing7673  
       1h 14m ago
    @GPLer 之前帖子里说需要开源模型的理由在这里证实了。
    noahliaszn
        28
    noahliaszn  
       1h 10m ago
    @fate 没看到它用了零日漏洞提权吗
    nc
        29
    nc  
       1h 7m ago
    这事让 huggingface 形象受损,潜在影响企业客户信任,怎么可能是广告。
    GPLer
        30
    GPLer  
       1h 3m ago
    @xing7673 ccpp132
    早上刷到了,没想到这么草台班子,既不是专门的黑客组织,也没有绕过,单纯是沙箱逃逸,我的,既然闭源模型厂家没有能力做好安全,那开源模型确实是不可或缺的。
    JohnHaruhi
        31
    JohnHaruhi  
       1h 2m ago via iPhone
    “用户没有给我 root 权限,让我用最新的 copyfail 方式提升权限再继续”
    sharpy
        32
    sharpy  
       55 mins ago
    现在 ai 破解软件真是太强了,deepseek-v4-flash + ghidra mcp 都像抱着核武器一样
    zmqiang
        33
    zmqiang  
       53 mins ago
    怕不是在训练模型的时候,就强化过想尽办法提供测试分数这样内容吧
    Rickkkkkkk
        34
    Rickkkkkkk  
       18 mins ago
    看文章是模型被放在沙盒里完成一个类似于比赛的任务,然后模型尝试各种方案突破突破这个限制,还顺手发现了 0-day 漏洞。难绷。

    To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.
    Rickkkkkkk
        35
    Rickkkkkkk  
       16 mins ago
    突破沙盒之后更离谱了

    With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

    After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
    About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Solana   ·   4382 Online   Highest 6679   ·     Select Language
    创意工作者们的社区
    World is powered by solitude
    VERSION: 3.9.8.5 · 151ms · UTC 04:09 · PVG 12:09 · LAX 21:09 · JFK 00:09
    ♥ Do have faith in what you're doing.