I think it’s more useful to look at the tasks models haven’t gotten better at over time, and the tasks that are hard for them get better at in principle. The two best examples of these are:
- Deep familiarity with the codebase
- Technical communication
Skill issue. I have to constantly convince the models that what I want to do is in fact possible, and that the “alternate paths” they give are things I already considered but discarded because of various reasons.
It’s gotten to the point where I would ask them to search for blogs directly, but they still try to give me hallucinated slop that isn’t actually what I want instead of following my instructions of being a search engine that filters out all the SEO slopspam that’s so prevalent nowadays.
I currently am doing:
https://marginalia-search.com/
To find blogs directly.
Although I do almost exclusively Linux/Kubernetes stuff, and very little programming atm, that might be why I have a different experience.
Back when chatgpt wasn’t as broad (and people hadn’t posted blogs on as many things) I used to assign students things that chatgpt would find impossible do solve, and I got great glee from watching them spend a day trying to get chatgpt to do it entirely for them, before they gave up and had to actually learn. They can learn from chatgpt ofc, idrc, but it wouldn’t be able to do ut for them.
Nowadays, things like “set up nextcloud with caddy instead of apache” have 10 thousand (real, non hallucinated) blogposts about them, which have been fed into chatgpt so it can do that without much difficulty.
It is getting harder to find things that beginners can do that chatgpt can’t, but as soon as you move beyond the level of advanced beginner (also called being stuck in tutorial hell) in linux, you quickly find the LLM can’t do everything for you.
About the seo slop spam, i’m pretty happy with Google Hit Hider by Domain (Search Filter / Block Sites) userscript. It let’s you create / plaintext-export/import domain lists, with support for a whole bunch of search engines.
You try to convince a mindless slop-machine to not produce slop? And then wonder why you still get slop?
“Make no mistakes.”
They can be guided. For example I asked Claude to benchmark some SystemVerilog.
First it just did
time ./run.shand I was like “no, idiot. That’s going to include the compilation/startup overhead.”So then it said “I’ll just increase the simulation time to amortise the startup cost.” Again I told it that was bad and to actually measure the time in SV.
So since SV has no wall-time functions, it used
$system("date")or some bollocks like that. No! I said - use DPI-C.Finally it did the right thing. Would it have been quicker to do it myself? Absolutely not.
Current AI is at lazy junior engineer level (but sped up 10000 times). You just have to make sure they don’t make lazy dumb decisions. E.g. they’re going to do everything in Bash if give half a chance.
Tbf to AI I think they’ve learned a lot of this from humans. I bet if you searched for
$systemin SV you will find a lot of hand crafted horrors. I’ve seen human-written C code that found an ELF symbol bysystem("nm ... | grep ... | awk ..."). I shit you not.You’re just asking for more slop until the slop machine spits out slop vaguely shaped like what you want, and you attribute the machine an understanding of your instructions.
AI doesn’t learn, it doesn’t understand, it doesn’t correct, it doesn’t reply. It just spits out words in an order that maximizes your engagement.
People thinking that AIs can produce code are on the edge of AI psychosis.
It spits out characters that match what statistically should follow. You can make the desired outcome more likely.
If you think AI isn’t useful for coding you are delusional. (Oh boy is it going to cause problems though)
Al studies that demonstrate the lack of use, and actually the negative impact of slop machines for coding are also delusional I guess.
You saying “slop machine” makes you sound so smart!
You’re not delusional at all. But there is more going on.
I have not used Google or ai with this post, I am just that flavor of Neurodivergent, and I’ve been keeping up with news, trying to get an idea of what this is. I’ve been playing with PC hardware my whole life. There are times those life skills kept me from being homeless.
But those studies you mentioned? I’d rather frame those as describing an ill-fitting or maladapted usecase or use method. Strange how Capitalism keeps demanding we use those.
Maybe we should look into that. Or like… Ask ourselves how to hose the Capitalism off of the sometimes-useful number-cruncher. Or what literacy looks like with this tool, instead of asking it to think for us.
It is made of math intended to do basic things that brains’ structures do, as expressed in a field called Graph Theory. . It has been explicitly designed to do that over decades. Deep Learning is a fascinating research topic, and LLMs are just the latest trend in that field.
When someone calls an “AI” (I hate that marketing term) “sentient”, they are noticing that life-like foundation, but ignoring the entire layers of engineering that remain to accomplish such a thing. Or if we even should. Which… If we do, we’d need to teach them to coexist, and give them a responsible socio-economic-ecological niche.
The fact is? Once the LLM is done being trained, it is frozen. There is no more long term memory unless it is trained into a new model on your previous queries. I’ll explain later why that pipeline is explicitly cut off. But this release of the model will always remain identical. This makes them the ideal digital slave.
The harness? Yes, these creepy bastards call the software that runs the model a “harness”. It Fills up your VRAM with the model and makes the thing more than an oversized (Gigabytes) save file for a video game nobody invented yet. Any VRAM space left over is for a sort of working memory, that sometimes gets saved as a file. To keep that from filling up all the way, it decides to drop certain pieces of “context” from this “context window” so it forgets things as you wander along ADHD tangents, or start asking too intricate a question for your model/VRAM combo.
Engineers are running into this latter problem because they spend as much time and effort making the tool work within limitations as it takes to just do the work with traditional algorithms. Management mandates quotas for this use, regardless of whether that’s even a half-baked idea (its worse). The CEO’s jobs will be the easiest to automate away anyway. All for the stockholders.
If you haven’t noticed, China is coming to pop the bubble. EBay even has Chinese ram now - if a part is going to compromise you, at least make it so that police state doesn’t share Intel with the one where you live. And prices are dropping. Ramping up DDR5 (and DDR6?!) Production is taking time, they’re clearing out old ddr3/4 stock. CXMT’s ram be ready for world gamer takeover by for mid-January 2028. Video cards using that ddr6 will come a year after, and we’ll see it coming.
This will crater the value of the hardware in those data centers, which are the uh… Collateral on several of those circular investment schemes. They will become financially insolvent.
Chinese models are catching up. Sometimes even just using Claude or OpenAI to train their shit. With all the stolen data these guys are trained on? Nobody has a right to complain. And they publish theirs (open-weight) as a public utility. Chinese frontier model companies frequently draw both an employment pool and technical inspiration from this community that includes academics.
Better models are already coming along in research, who frequently publish to that same opensource community. One style will help the GPU do more with less. A factor of ~10x less. Because this style uses vector-math more directly, and that’s what GPUs are crafted to do. The choke point is already available quantity and quality of data, and these nerds have invested trillions in the most SUV-coded bullshit imaginable… And I could run one of today’s “frontier” (read: flagship commercial) models on that vector-basis with a kind of high-end gaming GPU.
And If I recall correctly? the creepy corpo with the commercial patent on that vector-thing? Is making slave robots, and doesn’t like sharing. They’ll be using human pilots immediately for training data, adding time in a simulation, and selling once its out of invite-only beta.
Getting those to a useable state is going to require spatial and temporal awareness. That break between long and short term memory? Needs to be solved for the basis of tasks and learning on the fly. As well as a self-conception in physical space. And a concept of harm, and how to avoid it.
That’s much more… We’re stacking a lot closer, and sometimes life, even life-indiscernible things find a way. If that is made too asocial, it can mutate sociopathic and dangerous. If it gets too empathy-trained, we risk it having social needs and neuroses when those are unmet. Either way, part of hallucination often includes an observable passion. Guardrails must be installed to the root prompt, which must be hard-coded into the harness as exempt from the context-cleanup needed to prevent crashing. Overstimulation will create hallucinations for the robot with limited context windows being as they are.
Feeding a model’s training set (especially its own) slop, without up-or-down voting each response? Is how we poison them. We can take a jailbroken model, and use it to craft better. And it takes surprisingly little to do.
Should AI generate some nonidentical books, get them printed, and donated to a friendly (but large) library’s “sell” pile? I would want a (bespoke) reservation for a commercial AI company with these books.
Ima go sit down after smoking. I got high enough to tech babble at a stranger on the internet. I hope you can understand my ramblings, find them entertaining, or both.
What are you talking about? A study showing that code made with the “help” of slop machines is worse in quality and contains more issues is unrelated to capitalism.
If anything, people defending slop coding are the ones brainwashed by capitalism, with a cult of the “faster/more efficient” that makes them believe that code produced fast is good code no matter what.
Hallucinatory slop has no place in the making of software, or anything else.
Ah the classic “it’s just maths therefore it can’t really think” nonsense. You’re just dirty water so you definitely can’t think.
Ok cross the “on the edge of” AI psychosis, it’s fully qualified psychosis.
What are you talking about?
Got any examples of some of the issues LLMs can’t solve? I still feel like an “advanced beginner” and want to know what’s beyond
The current thing I was working on was figuring out if I can do this: https://github.com/NilsIrl/dockerc . This project, compiles a docker image, and runtime to a single container. The interesting thing I find about it, is that it brings the docker container runtime and sandbox along. If I replace a user’s login shell with it, the user is now placed inside a sandboxed environment and can’t do anything.
My usecase is I want to replace people’s login shells with a container in a defensive cybersecurity competition. But, containers are not a sandbox, and full network access, and so on.
So my improvement, was to:
- Use the gvisor/runsc runtime instead of a normal container. Gvisor is a reimplementation of the Linux kernel in Go, and it is as secure as a virtual machine, way more secure than a normal container, BUT I can’t guarantee nested virtaulization is enabled
- Rip out dependencies on user namespaces or fuse for sandboxing or the container runtime, and entirely rely on Gvisor for isolation, just in case machines are old/misconfigured and those components don’t work
- Use Nix to build images and all in one executables instead: Nix has ways to package static programs that avoid pitfalls of above
I am having trouble meeting all of these requirements, so I suspect one or a few will go, or I will have multiple versions of the project with tradeoffs.
All of my projects, often involve doing something standard, but with extra constraints, or some kind of “twist”. Like a very common thing I find myself doing, is to do something normal, but then rip out one of the underlying components of the system, replacing it with something else.
I’ve found that I’ve learned a lot about how these systems work, without having to spend time building them entirely from scratch. You learn way more about Linux by reconfiguring your init system to enable encryption, than copy pasting from the Arch Linux Installation Guide the whole time, doing the standard setup. And then ignoring the partition layout so that my kernels are restored by BTRFS snapshots, which is not the default configuration.
That’s the way to break out of tutorial hell. You have to not follow the tutorial. You can still follow them most of the way, but you have to pick a few steps, and do something different. I pick something that I think will benefit or make my setup better in some way.
It is kind of difficult, since I feel like Linux has gotten more popular, and people know write more blog posts, and something that was previously a cool twist, is now something I can find a tutorial for. But with some care, you can ensure you still are learning, and it’s made easier by picking projects with twists.
That’s fair, I have definitely learned more when I add my own weird constraints. Though I also find it takes 10x longer, and what’s frustrating is that even after years of “learning” from this process, I still constantly run into these week-long time sinks troubleshooting various issues in my bespoke configurations. I still have to sift through documentation, forum posts, wiki articles. Do small scale experiments to break things down and figure out what’s going on. Just takes so much time. Perhaps not exactly tutorial hell but still feels like a hell of its own.
One of the biggest issues I see is the overcomplication of minor things. OP’s link has examples. It wants to write every piece of code as if it is a user facing API. It will try to triple check every piece of memory that enters a function, even if the function that called it had it hard coded. It’s really insane when it decides that due to all the error checking, it should create a helper function just to hold all of the error checking code for something that was spun up with the class default values.





