<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Agent Ecosystem Glossary</title><link>https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/</link><description>Recent content in Evaluation on Agent Ecosystem Glossary</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Benchmarks</title><link>https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/benchmarks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/benchmarks/</guid><description>&lt;h1 id="benchmarks"&gt;Benchmarks&lt;a class="anchor" href="#benchmarks"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;Standardized datasets and tasks used to measure and compare LLM and/or agent&#10;performance across capabilities.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="arc"&gt;ARC&lt;a class="anchor" href="#arc"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Definition&lt;/strong&gt;: acronym for &lt;em&gt;AI2 Reasoning Challenge&lt;/em&gt;; benchmark measuring question answering and&#10;reasoning through more than 7,000 grade-school natural science questions&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Purpose&lt;/strong&gt;: evaluates an LLM&amp;rsquo;s ability to reason over knowledge, includes an easy set and a&#10;challenge set of harder questions requiring multi-step reasoning&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Related Terms&lt;/strong&gt;: &lt;a href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/#benchmark"&gt;benchmark&lt;/a&gt;, &lt;a href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/#evaluation-1"&gt;evaluation&lt;/a&gt;, &lt;a href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/#few-shot"&gt;few-shot&lt;/a&gt;, &lt;a href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/benchmarks/#mmlu"&gt;MMLU&lt;/a&gt;, &lt;a href="https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/#zero-shot"&gt;zero-shot&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Source&lt;/strong&gt;: &lt;a href="https://arxiv.org/abs/1803.05457"&gt;arXiv: &amp;ldquo;Think you have Solved Question Answering? Try ARC&amp;rdquo; by Clark et al.&lt;/a&gt;&lt;/p&gt;</description></item><item><title>Metrics &amp; Scoring</title><link>https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/metrics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://rhyannonjoy.github.io/agent-ecosystem-glossary/evaluation/metrics/</guid><description>&lt;h1 id="metrics--scoring"&gt;Metrics &amp;amp; Scoring&lt;a class="anchor" href="#metrics--scoring"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;Approaches used to quantify LLM and/or agent performance across the&#10;benchmarks and customized evaluations.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="accuracy"&gt;accuracy&lt;a class="anchor" href="#accuracy"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Definition&lt;/strong&gt;: also known as precision; percentage of correct predictions made by an LLM; foundational&#10;to evaluation, the most widely reported scoring metric in benchmarks and/or leaderboards; paired with&#10;recall and combined into the F1 score&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Purpose&lt;/strong&gt;: serves as the primary quantitative metric across classification; provides a single number&#10;for comparing how often an LLM produces a correct answer&lt;/p&gt;</description></item></channel></rss>