<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	>
<channel>
	<title>Comments on: LLMs, the useful parts</title>
	<atom:link href="http://thetarpit.org/2026/llms-the-useful-parts/feed" rel="self" type="application/rss+xml" />
	<link>http://thetarpit.org/2026/llms-the-useful-parts</link>
	<description>"Now I feel like I know less about what that blog is about than I did before."</description>
	<pubDate>Mon, 17 Aug 2026 07:05:16 +0000</pubDate>
	<generator>http://thetarpit.org</generator>
	<sy:updatePeriod>hourly</sy:updatePeriod>
	<sy:updateFrequency>1</sy:updateFrequency>
		<item>
		<title>By: spyked</title>
		<link>http://thetarpit.org/2026/llms-the-useful-parts#comment-7652</link>
		<dc:creator>spyked</dc:creator>
		<pubDate>Thu, 30 Jul 2026 18:07:00 +0000</pubDate>
		<guid isPermaLink="false">http://thetarpit.org/?p=593#comment-7652</guid>
		<description>&lt;blockquote&gt;a large (tera/petabyte-sized) corpus of so-called "training data"; think: literature, academic databases, videos, emails, Facebook comments, in other words, all the data that a large "service provider" such as Google or Meta could have gathered in a few decades;&lt;/blockquote&gt;

One of the problems with this approach -- which is quite obvious, although it might not be obvious to you, so let me spell it out clearly over here -- is that, in attempting to be an efficient mind-reader, the LLM agent will often assume the wrong thing. This, by the way, is in my humble opinion one of the most life-like characteristics of LLMs.

Take any life form -- say, a grapevine. Its characteristics are a function of both genetics, in the potential they provide via reproduction, and of the environment. Thus a quality or another of the grapes it produces will stem both from its heritage *and* the soil it inhabits, the altitude where it grew, the wind, sun and so on and so forth. And it works the same with intelligent life forms: say, if you take a bunch of &lt;a href="http://thetarpit.org/2018/lacul-morii?b=University&#038;e=Bucharest#select" rel="nofollow"&gt;UPB&lt;/a&gt; graduates, there's a high statistical likelihood that they will bear similar linguistic characteristics, which, for example, will help a small group of graduates who joined the local IT company communicate quite efficiently.

On the flip side, that's the problem with language, i.e. that it's largely contextual: if you were to bring in someone from Cluj, Iași, or even from Bucharest's own Universitate, they would have some trouble fitting in, even though supposedly engineering language works the same everywhere in the world (yeah, right). So, getting back to LLMs, they only exacerbate this problem, because of their huge training corpus.

In practice, this means that the user will have to go through great lengths to specify the problem context, but try as they might, that context cannot possibly be *complete*, due to the sheer fact that natural language is based on implicit context by its very... well, nature. So then the LLM agent will fill in the contextual gaps with whatever's more statistically likely, as derived from its training corpus. It often happens that that guessing will indeed resemble some form of mind-reading; but the flip side to that is that when it doesn't, it can lead to particularly nasty failure modes. Which makes the LLM the perfect programmable slot machine.</description>
		<content:encoded><![CDATA[<blockquote><p>a large (tera/petabyte-sized) corpus of so-called "training data"; think: literature, academic databases, videos, emails, Facebook comments, in other words, all the data that a large "service provider" such as Google or Meta could have gathered in a few decades;</p></blockquote>
<p>One of the problems with this approach -- which is quite obvious, although it might not be obvious to you, so let me spell it out clearly over here -- is that, in attempting to be an efficient mind-reader, the LLM agent will often assume the wrong thing. This, by the way, is in my humble opinion one of the most life-like characteristics of LLMs.</p>
<p>Take any life form -- say, a grapevine. Its characteristics are a function of both genetics, in the potential they provide via reproduction, and of the environment. Thus a quality or another of the grapes it produces will stem both from its heritage *and* the soil it inhabits, the altitude where it grew, the wind, sun and so on and so forth. And it works the same with intelligent life forms: say, if you take a bunch of <a href="http://thetarpit.org/2018/lacul-morii?b=University&#038;e=Bucharest#select" rel="nofollow">UPB</a> graduates, there's a high statistical likelihood that they will bear similar linguistic characteristics, which, for example, will help a small group of graduates who joined the local IT company communicate quite efficiently.</p>
<p>On the flip side, that's the problem with language, i.e. that it's largely contextual: if you were to bring in someone from Cluj, Iași, or even from Bucharest's own Universitate, they would have some trouble fitting in, even though supposedly engineering language works the same everywhere in the world (yeah, right). So, getting back to LLMs, they only exacerbate this problem, because of their huge training corpus.</p>
<p>In practice, this means that the user will have to go through great lengths to specify the problem context, but try as they might, that context cannot possibly be *complete*, due to the sheer fact that natural language is based on implicit context by its very... well, nature. So then the LLM agent will fill in the contextual gaps with whatever's more statistically likely, as derived from its training corpus. It often happens that that guessing will indeed resemble some form of mind-reading; but the flip side to that is that when it doesn't, it can lead to particularly nasty failure modes. Which makes the LLM the perfect programmable slot machine.</p>
]]></content:encoded>
	</item>
</channel>
</rss>
