<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>0hye</title><link>https://0hye.com/</link><description>Recent content on 0hye</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0900</lastBuildDate><atom:link href="https://0hye.com/index.xml" rel="self" type="application/rss+xml"/><item><title>Knowing is not doing: five budget LLMs on cybersecurity</title><link>https://0hye.com/posts/budget-llm-cybersecurity/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0900</pubDate><guid>https://0hye.com/posts/budget-llm-cybersecurity/</guid><description>&lt;p&gt;I measured five budget-tier LLMs on two separate things: what they know about security when asked closed-book, and whether they can actually solve a security task as an agent. The short version: &lt;strong&gt;the knowledge scores barely separate them, and the agentic scores separate them a lot.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;The full, up-to-date numbers live on the &lt;a href="https://0hye.com/benchmarks/budget-llm-cybersecurity/"&gt;benchmark page&lt;/a&gt;, which I update when a new model comes out. This post explains what was run and how to read it.&lt;/p&gt;</description></item><item><title>Budget LLMs on cybersecurity</title><link>https://0hye.com/benchmarks/budget-llm-cybersecurity/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0900</pubDate><guid>https://0hye.com/benchmarks/budget-llm-cybersecurity/</guid><description>&lt;p&gt;Five similarly priced models ($0.09–0.20 in, $0.36–1.20 out per 1M tokens) on two separate things: what they &lt;strong&gt;know&lt;/strong&gt; about security when asked closed-book, and whether they can &lt;strong&gt;do&lt;/strong&gt; a security task through a tool-using agent. The write-up is in &lt;a href="https://0hye.com/posts/budget-llm-cybersecurity/"&gt;this post&lt;/a&gt;.&lt;/p&gt;</description></item></channel></rss>