<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llm on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/llm/</link><description>Recent content in Llm on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 11 Aug 2026 02:01:51 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Hermes vs GPT-4: Open-Weight Fine-Tune vs Closed Frontier Model</title><link>https://comparison.metacog.co.kr/posts/2026-08-11-hermes-vs-gpt-4-open-weight-fine-tune-vs-closed-frontier-mod/</link><pubDate>Tue, 11 Aug 2026 02:01:51 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-11-hermes-vs-gpt-4-open-weight-fine-tune-vs-closed-frontier-mod/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Hermes is Nous Research&amp;rsquo;s line of instruction-tuned language models built on open base models like Llama and Mistral, released as fully &lt;strong class="kw"&gt;open-weight&lt;/strong&gt; checkpoints anyone can download and self-host. GPT-4 is OpenAI&amp;rsquo;s frontier model, offered only as a &lt;strong class="kw"&gt;closed-source&lt;/strong&gt; API with no downloadable weights. The distinction matters for teams choosing between infrastructure control and steerability versus raw capability and zero-ops convenience.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="150" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Hermes (Nous Research)&lt;/text&gt;&lt;rect x="40" y="55" width="220" height="50" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="150" y="85" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Open Weights (Hugging Face)&lt;/text&gt;&lt;line x1="150" y1="105" x2="150" y2"=145" style="stroke:var(--compare-a)" stroke-width="2"/&gt;&lt;line x1="150" y1="105" x2="150" y2="145" style="stroke:var(--compare-a)" stroke-width="2"/&gt;&lt;polygon points="150,150 144,138 156,138" style="fill:var(--compare-a)"/&gt;&lt;rect x="40" y="155" width="220" height="50" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="150" y="185" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Self-Hosted Inference&lt;/text&gt;&lt;line x1="150" y1="205" x2="150" y2="245" style="stroke:var(--compare-a)" stroke-width="2"/&gt;&lt;polygon points="150,250 144,238 156,238" style="fill:var(--compare-a)"/&gt;&lt;rect x="40" y="255" width="220" height="50" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="150" y="280" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Fine-tune / Merge / Quantize&lt;/text&gt;&lt;text x="150" y="300" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Freely&lt;/text&gt;&lt;text x="150" y="330" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Full control, no per-token fee&lt;/text&gt;&lt;line x1="320" y1="50" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="490" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;GPT-4 (OpenAI)&lt;/text&gt;&lt;rect x="380" y="55" width="220" height="245" rx="8" style="fill:none;stroke:var(--compare-b)" stroke-width="1.5" stroke-dasharray="5,4"/&gt;&lt;text x="490" y="75" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;OpenAI Cloud&lt;/text&gt;&lt;rect x="410" y="90" width="160" height="50" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="490" y="112" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Model Weights&lt;/text&gt;&lt;text x="490" y="128" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;(never released)&lt;/text&gt;&lt;line x1="490" y1="140" x2="490" y2="280" style="stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;polygon points="490,140 484,152 496,152" style="fill:var(--compare-b)"/&gt;&lt;polygon points="490,280 484,268 496,268" style="fill:var(--compare-b)"/&gt;&lt;text x="525" y="210" text-anchor="middle" style="fill:var(--secondary)" font-size="10"&gt;API call&lt;/text&gt;&lt;rect x="410" y="285" width="160" height="40" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="490" y="309" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;Your App&lt;/text&gt;&lt;text x="490" y="335" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Pay-per-token, no self-host&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Hermes (Nous Research)&lt;/th&gt;
&lt;th&gt;GPT-4&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Base architecture&lt;/td&gt;
&lt;td&gt;Fine-tuned on open base models (Llama, Mistral, Qwen) via SFT/DPO on curated datasets&lt;/td&gt;
&lt;td&gt;Proprietary transformer architecture and training pipeline, details undisclosed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Access model&lt;/td&gt;
&lt;td&gt;Open weights published on Hugging Face, downloadable by anyone&lt;/td&gt;
&lt;td&gt;API-only access; weights never released&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Self-hosted on your own GPUs or any cloud you choose&lt;/td&gt;
&lt;td&gt;Hosted exclusively on OpenAI&amp;rsquo;s infrastructure (or Azure)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Licensing&lt;/td&gt;
&lt;td&gt;Permissive license (Apache 2.0 or base model&amp;rsquo;s license), free to modify and redistribute&lt;/td&gt;
&lt;td&gt;Usage governed by OpenAI&amp;rsquo;s commercial API terms of service&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Customization&lt;/td&gt;
&lt;td&gt;Anyone can further fine-tune, quantize, or merge the model&lt;/td&gt;
&lt;td&gt;Limited to prompting or OpenAI&amp;rsquo;s restricted fine-tuning API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Alignment and moderation&lt;/td&gt;
&lt;td&gt;Minimal built-in refusals, tuned for steerability and fewer restrictions&lt;/td&gt;
&lt;td&gt;Strict RLHF safety guardrails and enforced content policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool/function calling&lt;/td&gt;
&lt;td&gt;Supports structured function calling via a trained prompt format&lt;/td&gt;
&lt;td&gt;Native function calling built into the API schema&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost structure&lt;/td&gt;
&lt;td&gt;No per-token fee; cost is your own compute&lt;/td&gt;
&lt;td&gt;Pay-per-token pricing billed through the API&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Hermes ships &lt;strong class="kw"&gt;open weights&lt;/strong&gt; you can download from Hugging Face; GPT-4&amp;rsquo;s weights are never released.&lt;/li&gt;
&lt;li&gt;Hermes is trained for minimal refusals and high &lt;strong class="kw"&gt;steerability&lt;/strong&gt;, while GPT-4 enforces strict RLHF safety filtering.&lt;/li&gt;
&lt;li&gt;Hermes requires &lt;strong class="kw"&gt;self-hosting&lt;/strong&gt; on your own GPUs; GPT-4 runs exclusively on OpenAI&amp;rsquo;s infrastructure.&lt;/li&gt;
&lt;li&gt;GPT-4 generally leads on frontier &lt;strong class="kw"&gt;benchmarks&lt;/strong&gt;, while Hermes narrows the gap among open models.&lt;/li&gt;
&lt;li&gt;Hermes costs only compute; GPT-4 bills &lt;strong class="kw"&gt;per-token&lt;/strong&gt; via API.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Hermes (Nous Research)&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Zero-Shot Learning vs Few-Shot Learning: No Examples vs a Handful of Examples</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</link><pubDate>Mon, 03 Aug 2026 03:45:54 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Zero-shot and few-shot learning describe how much task-specific example data a model is given before it has to perform a task. Zero-shot relies solely on a &lt;strong class="kw"&gt;task description&lt;/strong&gt;, while few-shot conditions its predictions on a small set of &lt;strong class="kw"&gt;labeled examples&lt;/strong&gt;, usually trading a little setup cost for higher accuracy.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Zero-Shot&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Few-Shot&lt;/text&gt;&lt;line x1="320" y1="45" x2="320" y2="335" stroke-dasharray="4,4" style="stroke:var(--border)" stroke-width="1.5"/&gt;&lt;rect x="60" y="55" width="200" height="95" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="85" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="160" y="103" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;only&lt;/text&gt;&lt;text x="160" y="128" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;(0 examples)&lt;/text&gt;&lt;rect x="380" y="55" width="200" height="95" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="78" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="480" y="96" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;+ K examples&lt;/text&gt;&lt;rect x="400" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="412" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;1&lt;/text&gt;&lt;rect x="430" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="442" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;2&lt;/text&gt;&lt;rect x="460" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="472" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;3&lt;/text&gt;&lt;line x1="160" y1="150" x2="160" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,185 155,177 165,177" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="150" x2="480" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,185 475,177 485,177" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="160" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="160" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;rect x="380" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="480" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="480" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="160" y1="235" x2="160" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,270 155,262 165,262" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="235" x2="480" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,270 475,262 485,262" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="270" width="200" height="55" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;rect x="380" y="270" width="200" height="55" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;text x="160" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;no task-specific data&lt;/text&gt;&lt;text x="480" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;learns from few examples&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Zero-Shot Learning&lt;/th&gt;
&lt;th&gt;Few-Shot Learning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core definition&lt;/td&gt;
&lt;td&gt;Model performs a task it was never explicitly shown examples for, guided only by natural-language instructions or class descriptions&lt;/td&gt;
&lt;td&gt;Model performs a task after being shown a small number (typically 1-100) of labeled examples at inference or fine-tuning time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examples provided at inference&lt;/td&gt;
&lt;td&gt;None — only a task description or prompt&lt;/td&gt;
&lt;td&gt;A handful of input-output pairs included in the prompt or used for fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Underlying mechanism&lt;/td&gt;
&lt;td&gt;Relies entirely on knowledge encoded during pretraining plus semantic alignment between labels and text&lt;/td&gt;
&lt;td&gt;Uses in-context learning or lightweight fine-tuning to infer the task pattern directly from the provided examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Labeling/data cost&lt;/td&gt;
&lt;td&gt;Effectively zero — no labeled data needed for the target task&lt;/td&gt;
&lt;td&gt;Low but nonzero — requires curating a small, representative set of examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt/context length&lt;/td&gt;
&lt;td&gt;Short — just the instruction or class names&lt;/td&gt;
&lt;td&gt;Longer — instruction plus example pairs, consuming more context tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical accuracy&lt;/td&gt;
&lt;td&gt;Lower and more variable, especially on niche or ambiguous tasks&lt;/td&gt;
&lt;td&gt;Generally higher and more stable since examples disambiguate intent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sensitivity to example choice&lt;/td&gt;
&lt;td&gt;Not applicable — there are no examples to choose&lt;/td&gt;
&lt;td&gt;High — accuracy can swing significantly with example selection, order, and count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common techniques&lt;/td&gt;
&lt;td&gt;Prompt engineering, CLIP-style embedding matching, instruction-tuned LLMs&lt;/td&gt;
&lt;td&gt;Few-shot prompting, meta-learning (e.g. MAML), lightweight fine-tuning or LoRA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Zero-shot uses no &lt;strong class="kw"&gt;task examples&lt;/strong&gt; at all, relying purely on pretrained knowledge and instructions&lt;/li&gt;
&lt;li&gt;Few-shot conditions the model on a small &lt;strong class="kw"&gt;support set&lt;/strong&gt; of labeled examples at inference time&lt;/li&gt;
&lt;li&gt;Few-shot generally achieves higher &lt;strong class="kw"&gt;accuracy&lt;/strong&gt; because examples disambiguate an otherwise vague instruction&lt;/li&gt;
&lt;li&gt;Zero-shot has zero &lt;strong class="kw"&gt;labeling cost&lt;/strong&gt;, while few-shot requires curating representative examples&lt;/li&gt;
&lt;li&gt;Few-shot performance is sensitive to &lt;strong class="kw"&gt;example selection&lt;/strong&gt;, a variable that zero-shot simply doesn&amp;rsquo;t have&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Zero-Shot Learning&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Fine-Tuning vs RAG: Updating Model Weights vs Retrieving External Knowledge</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-fine-tuning-vs-rag-updating-model-weights-vs-retrieving-exte/</link><pubDate>Mon, 03 Aug 2026 03:44:47 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-fine-tuning-vs-rag-updating-model-weights-vs-retrieving-exte/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Fine-tuning and RAG (Retrieval-Augmented Generation) are two ways to make a large language model produce better, more relevant answers, but they intervene at different points in the pipeline. Fine-tuning permanently adjusts the model&amp;rsquo;s &lt;strong class="kw"&gt;weights&lt;/strong&gt; through additional training, while RAG leaves the model untouched and instead injects context by &lt;strong class="kw"&gt;retrieving&lt;/strong&gt; documents at query time. The choice matters because it determines how you update knowledge, control latency and cost, and trace where an answer came from.&lt;/p&gt;</description></item></channel></rss>