comments (10)

  • I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!

    I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.

    leobuskin

  • Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/

    Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.

    I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?

    z4y5f3

  • This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.

    How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.

    aliljet

  • I might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it?

    It does feel like it respects both me and my time.

    Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.

    hypfer

  • > Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.

    I appreciate they don't just take the opportunity to self-glaze.

    aand16

  • Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it.

    Original images: https://image.non.io/neonRamenDesigns.webp

    GLM 5.3 build: https://html.non.io/neonRamenGLM5.3

    Opus 5 build for comparison: https://html.non.io/neonRamen

    For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.

    The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.

    jjcm

  • This will be roughly on pair with Kimi K3, but using a third of its parameters.

    Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.

    Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.

    bertili

  • > Scaling post-training is all we did for GLM-5.3.

    Love this opening line. And wow, great results.

    > As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.

    wxw

  • It's a smart model, but it has huge troubles staying on task... it constantly builds things I didn't ask it to.

    Might have its uses in a tightly controlled harness, but for coding I'm switching back to a mix of K3/GLM-5.2/Deepseek V4 Flash 0731

    kaeluka

  • OpenAI and Anthropic need to just go ahead and give people access to the cyber models.

    Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.

    virgildotcodes