<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software]]></title><description><![CDATA[<p>somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i/o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software</p>]]></description><link>https://forum.fedi.dk/topic/55b146ba-241e-4407-9260-73c90ad0a07e/somehow-i-only-recently-learned-that-llms-need-the-whole-conversation-fed-back-to-them-on-each-prompt-so-their-i-o-cost-scales-as-o-n-2-something-that-would-be-considered-completely-unacceptable-in-almost-any-other-production-network-accessible-software</link><generator>RSS for Node</generator><lastBuildDate>Fri, 24 Jul 2026 04:07:49 GMT</lastBuildDate><atom:link href="https://forum.fedi.dk/topic/55b146ba-241e-4407-9260-73c90ad0a07e.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 23 Jul 2026 08:35:04 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 10:31:06 GMT]]></title><description><![CDATA[<p><span><a href="/user/jcoglan%40mastodon.social">@<span>jcoglan</span></a></span> Sort of, but not really in reality. What happens is that you keep the context cached in VRAM and only transfer what's new.</p><p>It's like complaining on GIMP needing to have the whole image in RAM to be able to work on it.</p>]]></description><link>https://forum.fedi.dk/post/https://swecyb.com/users/troed/statuses/116968827576325386</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://swecyb.com/users/troed/statuses/116968827576325386</guid><dc:creator><![CDATA[troed@swecyb.com]]></dc:creator><pubDate>Thu, 23 Jul 2026 10:31:06 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:58:46 GMT]]></title><description><![CDATA[<span><a href="/user/schrotthaufen%40mastodon.social" rel="ugc">@<span>schrotthaufen</span></a></span> <span><a href="/user/jcoglan%40mastodon.social" rel="ugc">@<span>jcoglan</span></a></span> And coding agents gaslight the user by injecting "system reminders" prefixes that are invisible to the user but the LLM sees them as user messages and regularly refer back to those as user requirements.]]></description><link>https://forum.fedi.dk/post/https://infosec.place/objects/28553edb-40ef-4477-a6e3-930e6b68f391</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://infosec.place/objects/28553edb-40ef-4477-a6e3-930e6b68f391</guid><dc:creator><![CDATA[buherator@infosec.place]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:58:46 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:56:08 GMT]]></title><description><![CDATA[<p><span><a href="/user/jcoglan%40mastodon.social">@<span>jcoglan</span></a></span> Have-You-Tried-Turning-It-Off-And-On-Again as a service</p>]]></description><link>https://forum.fedi.dk/post/https://tech.lgbt/users/henrahmagix/statuses/116968454095877657</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://tech.lgbt/users/henrahmagix/statuses/116968454095877657</guid><dc:creator><![CDATA[henrahmagix@tech.lgbt]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:56:08 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:54:49 GMT]]></title><description><![CDATA[<p><span><a href="/user/jcoglan%40mastodon.social">@<span>jcoglan</span></a></span> <span><a href="/user/buherator%40infosec.place">@<span>buherator</span></a></span> This also means you can “gaslight” the LLM by manipulating the conversation before it’s sent back to the LLM</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/schrotthaufen/statuses/116968448924118318</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/schrotthaufen/statuses/116968448924118318</guid><dc:creator><![CDATA[schrotthaufen@mastodon.social]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:54:49 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:53:36 GMT]]></title><description><![CDATA[<p>they invented a technology business with super-linear marginal cost of production I am losing my *entire mind*</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968444153600555</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968444153600555</guid><dc:creator><![CDATA[jcoglan@mastodon.social]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:53:36 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:52:20 GMT]]></title><description><![CDATA[<p>SAMA ET AL: why can't we make a profit with this revolutionary technology<br />ME: your operating cost scales super-linearly with the value it delivers to users<br />SAMA ET AL: better tell the government we have invented the angels from evangelion or something</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968439217901540</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968439217901540</guid><dc:creator><![CDATA[jcoglan@mastodon.social]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:52:20 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:49:55 GMT]]></title><description><![CDATA[<p>this also makes it hard to take seriously claims that the marginal cost of inference is not a serious contributor to the resource use of data centres. your program has polynomial running cost and people run it for hours on end</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968429653967850</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968429653967850</guid><dc:creator><![CDATA[jcoglan@mastodon.social]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:49:55 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:47:55 GMT]]></title><description><![CDATA[<p>they ran into this and then didn't figure out a way to externalise the state in some fixed-size form so the LLM can resume from there. you just have to restart the computer and redo everything</p><p>in any other program this would be considered a DoS vector and have CVEs issued</p><p>gosh I wonder why their costs are completely unmanageable</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968421848302865</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/jcoglan/statuses/116968421848302865</guid><dc:creator><![CDATA[jcoglan@mastodon.social]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:47:55 GMT</pubDate></item><item><title><![CDATA[Reply to somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i&#x2F;o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software on Thu, 23 Jul 2026 08:44:21 GMT]]></title><description><![CDATA[<p><span><a href="/user/jcoglan%40mastodon.social">@<span>jcoglan</span></a></span> i was thinking about this, and i believe that simple inference is O(n^2), because it is matrix multiplication of input vector (input tokens). The situation with refeeding the whole context, which also has hiddeng growth, or even “agentic” approach - launching secondary queries by agents - makes the situation far worse.<br />My take is that the general strategy of “AI” companies is - in principle - throwing exponential complexity brute force on problems and showing off that sometimes it gives interesting result.</p>]]></description><link>https://forum.fedi.dk/post/https://hachyderm.io/users/prema/statuses/116968407800097279</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://hachyderm.io/users/prema/statuses/116968407800097279</guid><dc:creator><![CDATA[prema@hachyderm.io]]></dc:creator><pubDate>Thu, 23 Jul 2026 08:44:21 GMT</pubDate></item></channel></rss>