<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Benchmarks on 0hye</title><link>https://0hye.com/benchmarks/</link><description>Recent content in Benchmarks on 0hye</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 28 Sep 2026 00:00:00 +0900</lastBuildDate><atom:link href="https://0hye.com/benchmarks/index.xml" rel="self" type="application/rss+xml"/><item><title>Budget LLMs on cybersecurity</title><link>https://0hye.com/benchmarks/budget-llm-cybersecurity/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0900</pubDate><guid>https://0hye.com/benchmarks/budget-llm-cybersecurity/</guid><description>&lt;p&gt;Five similarly priced models ($0.09–0.20 in, $0.36–1.20 out per 1M tokens) on two separate things: what they &lt;strong&gt;know&lt;/strong&gt; about security when asked closed-book, and whether they can &lt;strong&gt;do&lt;/strong&gt; a security task through a tool-using agent. The write-up is in &lt;a href="https://0hye.com/posts/budget-llm-cybersecurity/"&gt;this post&lt;/a&gt;.&lt;/p&gt;</description></item></channel></rss>