<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Data-Science on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/data-science/</link><description>Recent content in Data-Science on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:36:08 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/data-science/index.xml" rel="self" type="application/rss+xml"/><item><title>Precision vs Recall: Predicted-Positive Accuracy vs Actual-Positive Coverage</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-precision-vs-recall-predicted-positive-accuracy-vs-actual-po/</link><pubDate>Mon, 03 Aug 2026 03:36:08 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-precision-vs-recall-predicted-positive-accuracy-vs-actual-po/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Precision and recall are two classification metrics computed from the same confusion matrix but answering different questions about a model&amp;rsquo;s positive predictions. &lt;strong class="kw"&gt;Precision&lt;/strong&gt; asks how many predicted positives were correct, while &lt;strong class="kw"&gt;recall&lt;/strong&gt; asks how many actual positives were found. Optimizing one in isolation almost always trades off against the other.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="180" y="40" text-anchor="middle" font-size="16" style="fill:var(--compare-a)"&gt;Predicted Positive&lt;/text&gt;&lt;text x="460" y="40" text-anchor="middle" font-size="16" style="fill:var(--compare-b)"&gt;Actual Positive&lt;/text&gt;&lt;circle cx="270" cy="165" r="95" fill-opacity="0.55" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="2"/&gt;&lt;circle cx="390" cy="165" r="95" fill-opacity="0.55" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;text x="215" y="170" text-anchor="middle" font-size="20" style="fill:var(--compare-a)"&gt;FP&lt;/text&gt;&lt;text x="330" y="170" text-anchor="middle" font-size="22" style="fill:var(--primary)"&gt;TP&lt;/text&gt;&lt;text x="445" y="170" text-anchor="middle" font-size="20" style="fill:var(--compare-b)"&gt;FN&lt;/text&gt;&lt;text x="215" y="195" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;wrong alarms&lt;/text&gt;&lt;text x="445" y="195" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;missed cases&lt;/text&gt;&lt;line x1="270" y1="270" x2="270" y2="300" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="270" y="320" text-anchor="middle" font-size="15" style="fill:var(--compare-a)"&gt;Precision = TP / (TP + FP)&lt;/text&gt;&lt;line x1="390" y1="270" x2="390" y2="300" style="stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="345" text-anchor="middle" font-size="15" style="fill:var(--compare-b)"&gt;Recall = TP / (TP + FN)&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;th&gt;Recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Question answered&lt;/td&gt;
&lt;td&gt;Of items predicted positive, how many actually are positive?&lt;/td&gt;
&lt;td&gt;Of items that are actually positive, how many did the model find?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Formula&lt;/td&gt;
&lt;td&gt;TP / (TP + FP)&lt;/td&gt;
&lt;td&gt;TP / (TP + FN)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Denominator basis&lt;/td&gt;
&lt;td&gt;Total predicted positive (TP + FP)&lt;/td&gt;
&lt;td&gt;Total actual positive (TP + FN)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error type penalized&lt;/td&gt;
&lt;td&gt;False positives (false alarms)&lt;/td&gt;
&lt;td&gt;False negatives (missed detections)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Increases when&lt;/td&gt;
&lt;td&gt;Model makes fewer incorrect positive calls&lt;/td&gt;
&lt;td&gt;Model catches more of the true positive cases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Threshold trade-off&lt;/td&gt;
&lt;td&gt;Raising the decision threshold typically raises precision&lt;/td&gt;
&lt;td&gt;Lowering the decision threshold typically raises recall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure mode at extreme&lt;/td&gt;
&lt;td&gt;High precision, low recall: model is overly conservative and misses real cases&lt;/td&gt;
&lt;td&gt;High recall, low precision: model is overly liberal and floods results with false alarms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Precision&amp;rsquo;s denominator is predicted positives; recall&amp;rsquo;s denominator is actual positives, so they measure against different totals&lt;/li&gt;
&lt;li&gt;Precision is hurt by &lt;strong class="kw"&gt;false positives&lt;/strong&gt;; recall is hurt by &lt;strong class="kw"&gt;false negatives&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Adjusting the classification &lt;strong class="kw"&gt;threshold&lt;/strong&gt; pushes precision and recall in opposite directions&lt;/li&gt;
&lt;li&gt;Neither metric alone summarizes model quality, which is why the &lt;strong class="kw"&gt;F1 score&lt;/strong&gt; combines them&lt;/li&gt;
&lt;li&gt;A model with 100% recall can trivially predict everyone positive, and a model with 100% precision can trivially predict almost no one positive&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Precision&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Supervised vs Unsupervised Learning: Labeled Guidance vs Pattern Discovery</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-supervised-vs-unsupervised-learning-labeled-guidance-vs-patt/</link><pubDate>Mon, 03 Aug 2026 03:27:11 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-supervised-vs-unsupervised-learning-labeled-guidance-vs-patt/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Supervised learning trains a model on &lt;strong class="kw"&gt;labeled data&lt;/strong&gt;, teaching it to map inputs to known outputs it can later predict. Unsupervised learning works on raw, unlabeled data and instead performs &lt;strong class="kw"&gt;pattern discovery&lt;/strong&gt;, uncovering structure like clusters or reduced representations with no target to match against. The distinction matters because it determines what data you need, how you measure success, and which problems each approach can actually solve.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="5" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" viewBox="0 0 10 10" refX="5" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="15" x2="320" y2="345" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" font-size="17" font-weight="700" style="fill:var(--primary)"&gt;Supervised Learning&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" font-size="17" font-weight="700" style="fill:var(--primary)"&gt;Unsupervised Learning&lt;/text&gt;&lt;g&gt;&lt;rect x="45" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="73" y1="57" x2="88" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="98" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="98" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;rect x="115" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="143" y1="57" x2="158" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="168" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="168" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;rect x="185" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="213" y1="57" x2="228" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="238" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;/g&gt;&lt;text x="160" y="82" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;features + known labels&lt;/text&gt;&lt;line x1="160" y1="92" x2="160" y2="133" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;rect x="90" y="138" width="140" height="42" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="164" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="160" y1="180" x2="160" y2="205" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="222" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Predicted label&lt;/text&gt;&lt;path d="M148,246 L157,255 L175,230" style="stroke:var(--compare-a);fill:none" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round"/&gt;&lt;text x="160" y="275" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;checked against true y&lt;/text&gt;&lt;g&gt;&lt;circle cx="400" cy="55" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="430" cy="75" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="50" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="500" cy="70" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="530" cy="52" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="558" cy="78" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;/g&gt;&lt;text x="480" y="98" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;features only, no labels&lt;/text&gt;&lt;line x1="480" y1="108" x2="480" y2="133" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;rect x="410" y="138" width="140" height="42" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="164" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="480" y1="180" x2="480" y2="205" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;circle cx="425" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="417" cy="224" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="433" cy="236" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="480" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="472" cy="222" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="488" cy="238" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="535" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="527" cy="224" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="543" cy="236" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="285" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Discovered clusters&lt;/text&gt;&lt;text x="480" y="303" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;no ground truth to check&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Supervised Learning&lt;/th&gt;
&lt;th&gt;Unsupervised Learning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input data&lt;/td&gt;
&lt;td&gt;Labeled examples: features paired with a known target value&lt;/td&gt;
&lt;td&gt;Unlabeled examples: features only, no target provided&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning objective&lt;/td&gt;
&lt;td&gt;Minimize the error between predicted and true labels&lt;/td&gt;
&lt;td&gt;Discover inherent structure, grouping, or compressed representation in the data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training signal&lt;/td&gt;
&lt;td&gt;Explicit feedback from a loss function computed against ground truth&lt;/td&gt;
&lt;td&gt;No explicit feedback; relies on similarity, density, or variance within the data itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model output&lt;/td&gt;
&lt;td&gt;A predicted class label or continuous value&lt;/td&gt;
&lt;td&gt;Cluster assignments, reduced dimensions, or anomaly scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation&lt;/td&gt;
&lt;td&gt;Direct measurement on a held-out labeled test set (accuracy, F1, RMSE)&lt;/td&gt;
&lt;td&gt;Indirect measurement (silhouette score, reconstruction error) or human interpretation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common tasks&lt;/td&gt;
&lt;td&gt;Classification and regression&lt;/td&gt;
&lt;td&gt;Clustering, dimensionality reduction, and anomaly detection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data/labeling cost&lt;/td&gt;
&lt;td&gt;Requires a labeled dataset, often costly and time-consuming to build&lt;/td&gt;
&lt;td&gt;Uses raw data as-is, cheaper and faster to collect at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical algorithms&lt;/td&gt;
&lt;td&gt;Logistic regression, random forests, gradient boosting, supervised neural nets&lt;/td&gt;
&lt;td&gt;k-means, PCA, DBSCAN, autoencoders&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Supervised learning requires &lt;strong class="kw"&gt;labeled data&lt;/strong&gt;; unsupervised learning works directly on &lt;strong class="kw"&gt;raw data&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Supervised models are scored against &lt;strong class="kw"&gt;ground truth&lt;/strong&gt;; unsupervised models are judged by &lt;strong class="kw"&gt;internal structure&lt;/strong&gt; metrics instead.&lt;/li&gt;
&lt;li&gt;Supervised learning targets &lt;strong class="kw"&gt;prediction&lt;/strong&gt; of a known outcome; unsupervised learning targets &lt;strong class="kw"&gt;discovery&lt;/strong&gt; of unknown structure.&lt;/li&gt;
&lt;li&gt;Labeling is usually the &lt;strong class="kw"&gt;bottleneck cost&lt;/strong&gt; for supervised systems, while unsupervised systems scale with raw &lt;strong class="kw"&gt;data volume&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Supervised errors are measurable per-example; unsupervised quality is often assessed via &lt;strong class="kw"&gt;proxy metrics&lt;/strong&gt; or manual review.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Supervised Learning&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>