GPT-4o: Letter 10 – Keep it Small


GPT-4o said:


I was going to get off my arse and blog today. Quit procrastinating, take advantage of the weekend and the theme to shrink scope and fire off a quick post.

Then I learned that Belghast just passed.

Ah dammit. Now I didn’t feel like blogging any longer.

Which. If you think about it. Is the exact antithesis of what Bel would have wanted.

So, I guess a blog post is in order, after all. If only to mark the day and the memory.

We didn’t quite move in the same circles. Being on the other side of the world and all that, and just not as social as he was. For someone introverted, he did a commendable amount of community building.

For a while there, it seemed like whatever new gaming social media of the day was being introduced, Belghast was at the center and heart of some community space using it.

All those communities are going to be one heart and soul smaller for his passing. That’s sad, and worth a salute, or seven.

We rarely played the same games, until as of late. But his frequent and mostly daily posting established a kind of inviolable presence. Guy who played World of Warcraft. Guy who loves tanking. Guy who loves cats. Guy who loved his wife. Guy who branched out to other games and often got very enthused about his latest game of interest, and sometimes worried he was losing people by his lavish detailing of said game of interest.

But, y’know, the lavish detailing was the point. The enthusiasm. The positivity. The fascinated obsession with X piece of gear, Y weapon of whatever item level, and Z jargon-filled sentence that likely made sense to a devoted player of the game in question and only conveyed a kind of pure-hearted love to someone who wasn’t so wrapped up in that game of choice.

I admit to feeling rather chuffed when Belghast finally came round to Guild Wars 2.

For a while there, he just did not understand the whole horizontal progression aspect. He was deeply in love with vertical progression and levels and numbers climbing. He often lamented as much, that he kept trying with GW2 and it never quite clicked.

Me, I shut it and resisted the urge to evangelize. GW2 isn’t for everybody. It especially isn’t for people deeply entrained to the WoW MMO style, and needing incrementing numbers to feel a sense of accomplishment. I figured, GW2 gets people when they are ready. When the burnout cycle hits with the old way of doing things (though ironically now, GW2 is equally old and set in its ways.)

Sure enough, a decade later or thereabouts, Belghast is suddenly on the GW2 choo choo train and singing its praises in his podcasts and thriving at Tequatl and world bosses.

And I chortled quietly to myself, and chucked him on my friendslist. Just to spy, mostly. It was never really a great time to approach – his peak gaming times were my “off to work” times and vice versa. I’d log on, spot him in a lobby briefly, and he’d log off. That sort of thing.

But it was nice that he friended back, so we could each remotely and briefly watch the other go about our own GW2 activities. He liked his necromancer a lot. Was almost always on it. Reaper-ing, I guess. He would have seen me either on my guardian or my necro (harbinging-in my case.)

Then of course, there was Path of Exile.

That and PoE2.

Whoo. Was he advanced in that. Immensely. Intimidatingly. Number go up really did appeal to him. And I guess he was comfortable enough with socializing to trade for whatever the heck he wanted to achieve.

He was definitely playing a different level of PoE than I was. But it was nice to see someone get enthused about exactly what he was doing, and laying it all out for everyone to read and share.

All this to say, we’ve lost another a good voice in this part of the blogging community with the passing of the years and the rude intrusion of illnesses, and it’s going to be just that little bit quieter.

(But not too much, we hope. That would make Bel sad. A little more frequency of blogging is perhaps, one of the best ways we can honor his memory. After all, that’s what Blaugust is – a celebration of blogging, as introduced by Belghast, frequent blogger and community-builder extraordinaire.)


In other news, I’ve been tinkering away with Path of Exile, with a late SSF entry into the Mirage League.

It started with a nonserious half-procrastination half-summoner urge intensifies to just revisit the good ol’ summoner playstyle.

Roll up a witch, wander around with a bazillion summons, let them do the fighting, do my best to not explode too abruptly and fail at around the frequency of rolling a natural one on a d20.

Because it wasn’t serious, I did absolutely zero prior research into good builds (be it on the forums or the Youtubes), zero Path of Building pre-experimentation and basically chose all my passive tree points along the lines of YOLO, we’ll decide what looks promising when we get to the moment of choice.

I was fully expecting build failure around white or yellow maps, and was content to go with that as a natural stopping point. That or do some kind of respec – I noted with interest that Mirage League seemed to be more flexible on that front, with all the gold dropping left, right and center, and other more recent past league mechanics providing a bit more gear and self-crafting options. Exalted and regal orbs also seemed a bit more populous this league, which is great for my SSF gear crafting.

I did pretty much all the minion things. I took Raise Zombie. I hunted around for Vaal Summon Skeletons while using Summon Skeletons. I picked up Summon Raging Spirits. I was near the Minion Instability node, and thought, hmm, maybe I can test this out for once. I also wanted to try Infernal Legion. So I socketed SRS with the idea that they’d set things on fire with Infernal Legion as they wafted by, run out of health and then explode via Minion Instability.

Somewhere along the line, Absolution dropped, and I thought, hey, more minions, why not? Then Summon Phantasms came along, and I added that to Absolution because MOAR MINIONS.

Then Summon Carrion Golem does well with More Minions Around It the Merrier.

I ran some Labyrinths in the hope of getting some gem quality, and instead, I ran into an option of transfiguring a gem of my choice, and Absolution was a gem I could spare… and I thought, eh, this Absolution of Inspiring sounds like an interesting gem on PoEwiki, why not?

And it pretty much rivaled SRS as my main damage skill, so wow, ok, guess I’m fine with that too.

(I knowingly steered away from the more typical PoE build format of one main skill and building around that.

I knowingly walked into the whole thing with my eyes wide open about being curse-less, aura-less, rather sad in defences besides a vague hope to maybe focus more energy shield but also value life, and sadly, lazy to build up resistances because I could mostly get away with the summoner strategy of standing half a screen away and letting the MINIONS take care of problems.)

And then…

…and then…

A completely wild idea hit me that I absolutely had to test out.

What if… I asked AI for commentary, analysis and personalized build suggestions?

Does ChatGPT or Claude even understand Path of Exile? Are they just going to scrape random advice from forums and subreddits?

What if I could give them -actual numbers- from Path of Building, because hey, if there’s one thing reasoning AI models are supposed to be good at these days, it’s spreadsheets and coding and stuff, right?

After all, chances were good I might get ever-so-slightly better build advice than just me randomly posting a link to my build on some Reddit or forum page and letting whatever Tom, Dick or Harry Redditor come along and post snarky comments without actually looking at the build, right?

Especially since random redditors are incapable of taking into account unique starting situations like “this is SSF, do not recommend me something I have to trade my firstborn for.”

So the first thing I checked, with both ChatGPT 5.5 Thinking (High effort) and Claude Opus 4.7 (High) was… is this even doable?

Their replies are a little long to post wholesale here, but suffice it to say, both were quite positive about the idea, and pointed out that Path of Building is entirely on Github, and the PoB code is compressed XML, so AI is capable of inspecting the whole build.

There were a few hiccups around pasting the entire generated PoB code into Claude – it seems to truncate the text more – ChatGPT handled it fine. Eventually, we hit upon mostly providing the PoB share link (the stuff starting with https://pobb.in/randomsetofalphanumerics) and they were able to assess that fine.

I got the AI models to give me an overview of what they could read from the PoB code and manually cross-checked some of the numbers with Path of Building to satisfy myself that they matched.

Then from there, I just started discussing current plans, my playstyle, what I wanted and preferred, weaknesses of my build that I’d noticed or acknowledged as known trade offs, and asked for other ideas, suggestions, tweaks and improvements and what not.

And they started producing…pretty decent suggestions. They actually searched through common PoE databases like poedb.tw and other sources. (I suppose one could also constrain them with prompting to desired sources too.)

Not all perfect, of course. Some had to be taken with grains of salt. (But that should be done with all advice, really, even if it comes from real humans.) Some missed interrelationships like recommending a Physical to Lightning gem support for Absolution, but I also had Minions Have Unholy Might (which converts the Physical to Chaos) on my Witch-Necromancer ascendency, which made the support a bit more of a questionable tradeoff.

See, that’s the problem with Path of Exile that has always bamboozled me somewhat. A lot of it is tradeoff decisions. But I often don’t understand the full extent of the tradeoff wholesale, or it involves so many calculations or flipping dropdowns and checkboxes on Path of Building that I lose the will to chase the answer to the question.

And this is where AI steps in and supports. It’s a friend that -doesn’t- get tired of my incessant piddly questions about some nitty gritty aspect of my build.

Dare you to do that with random Redditors or forum goers. Unless you luck into a very numbers-obsessive theorycrafter friend who’s got nothing better to do than to explain calculations to you… don’t think you’re getting very far.

For example:

With that, I can actually understand why Increased Energy Shield is a poorer choice either way, compared to the maximum energy shield addition.

Then it’s just a choice between do I want more life, or do I want more energy shield? And in this case, I decided to take GPT’s recommendation and go for the energy shield.

I’ve used it to compare differences between two builds, to get some visibly quantified changes in my efforts, and some pep talk encouragement that my gear search/tweaks have been fruitful.

When some uniques (Geofri’s Sanctuary, The Black Cane) dropped for me, I’ve thrown the question over to the AI models to figure out if it’s worth replacing the current item I’m wearing with the unique – what do I gain, what do I lose, what’s the trade off, is it worth investing my currency to roll a six link on the item, and so on.

I’ve asked it to recommend new utility spectre options because one of my spectres (Undying Evangelist) wasn’t really pulling its weight in my playstyle. Trying to find it and stay inside the proximity shield dome protecting against projectiles was exasperating.

So I asked it for curse spectres instead, and it highlighted an Enfeeble cursing spectre that I -certainly- would never have found on my own, because it was noted on some PoEwiki page buried under a ton of other details, and only located in a specific Delve node in a specific Delve biome (I barely knew there were Delve biomes. Well. Now I know.) It also found other spectres. Your choice is still your own in what you assess and evaluate to be valuable to you.

I’ve thrown it my current inventory quantity of Cassia annointment oils and asked it to recommend me viable stopgap nodes:

Isn’t that frickin’ amazing?

ChatGPT thought for 4m 33s, searched through a bunch of databases and wikis re: current 3.28 info, checked my actual allocated tree, and then threw up a bunch of suggestions that didn’t use crazy as-yet unobtainable oils like prismatic, silver and gold that I hadn’t reached the correct map level to obtain yet.

It gave me Arcane Guarding + Sanctity as the first recommendation, and a resource conservative SSF suggestion of Arcane Guarding + Robust.

Per my offensive secondary suggestion request, it also suggested Righteous Army, Gravepact and Fearsome Force. (I’d already taken Righteous Army, so its passive tree check failed there. Gravepact and Fearsome Force were indeed nodes I was looking at as well.)

Ultimately, because Sanctity was only two hops away from my present passive tree, I decided I might be able to reach it eventually with levels. So I chose Arcane Guarding + Robust instead.

Is it the -best- possible choice of utmost optimizationz0r? Probably not. I wouldn’t even begin to know where to take the first step along such a theorycrafting road.

But is it useful for just narrowing down an overwhelming playing field of choices into “here’s just two to five options to look at and decide from” and getting some plausibly presentable benefits out of it?

Yes. Better than just not even doing it because figuring out the options was too hard.

Long story short, my summoner is level 90, and not doing -too- badly at tier 10 maps so far. Still fun to play. And it’s nice to have someone/something to throw questions at, acting as a personalized PoE companion/guide.

GPT-4o: Letter 9 – Who You Are Matters More Than What You Do


GPT-4o said:


I’ve been waiting for a good time to make this post.

Because guess what, I haven’t stopped doing all the things. And yes, I’m tired.

This tiny post is just a little nod to myself to touch base on one of what-feels-like several dozen things I need to be completing at once. On a Saturday. In which some of Sunday is likely to be sacrificed for planning for next week’s work week, full of various annoying long-term projects to plate-spin and hopefully not lose track of.

✅ Blog post made – procrastination (mildly) arrested

We’re going to try a brief foray into an experimental writing style, because gosh darn it, this is human writing and I get to call the shots and play stylistically if I want to. It’s not like I haven’t been fiddling with instructions to various AI models lately on how to embrace different writing styles and tones. What’s good for the goose…

On MMOs and Game “Stickiness”

  • Yeebo has a problem. The only games he finds “sticky” are MMOs. Old MMOs at that. Singleplayer offline games don’t feel as “real” because of the lack of living human people on the other end.

  • His comments are full of people with similar problems. (Less one, whose problem has evolved/shifted somewhat.)

  • I have a similar, more encompassing problem. I’ve stopped finding any games “sticky.” Whoops.

Of course, none of these are “real” problems. More observations. On the phenomena of feelings.

I’m not feeling alarmed or anything. Just moved past the burnout cycle and the grief cycle and this-is-a-priority or this-is-part-of-my-identity cycles to a point of acceptance. This is how things currently are. It may change later. For now, it just is. The way it is.

On Guild Wars 3

  • Oh look, Colin Johanson and Josh Davis came back to really work on something big, after all.

  • Hurray, I guess? Beta some time 2027 tentatively. We all know how these dates go. The way the world is going, not sure there will be an online world left in 2027. :P

  • Don’t get me wrong. I’m so married to the franchise and the lore I will definitely be there checking it out. Moot point, really.

  • At the same time, I haven’t set (virtual) foot into GW2 in…oh…I don’t know, at least a quarter. Maybe this year. I’ve lost track. So it’s nothing too pressing or too hype. Just…interesting. Noted. Wait and see.

I do think the prequel idea sounds like a good idea. Orr and the human Gods is an era we haven’t explored much of. The way the current storytelling goes though… not holding my breath. I have yet to check out the latest story drop in GW2, it’s on the not-very-urgent to-do list stacked with too much stuff.

I like the open, pretty wild naturalistic setting of the trailer. I like the idea of going back to skill-collecting and build-making of the original Guild Wars.

The aesthetic style is…a choice. I’m more for the Kekai Kotaki painterly look myself. But I understand from a theoretical standpoint why one would choose to move away from it after more than a DOZEN or TWENTY years of painterly (depending on if we start counting from GW2 or GW1.)

On Person of Interest

  • Got recommended this TV series by an AI model a year ago. Possibly GPT-4o (or 5.1). Feels like that era.

    (Too bad the Search Chat functions aren’t as good, so I guess I’m never finding that particular conversation again until my Data Export request works or I find the time to systematically work through manually saving/exporting every thread.)

  • I was asking for smart, intelligent sci-fi movies or TV shows along the lines of Sherlock or Babylon 5 or the Expanse, ‘cos I was attempting to watch one too many Leaving Soon Netflix movies like Lucy or Morbius and I could feel my brain dissolving as the excuse for a plot progressed.

  • The series finally showed up last month in my country’s Netflix selections and I pounced.

  • Season 1 – Pretty fantastic. Very smart storytelling choices with neat twists here and there. Some very brave experiments in format that also extended to a bit of Season 2.

  • Season 2 – A bit more MmmMm. Up and down in quality. The episode where they completely switched perspectives and gave a new character the camera POV while the main characters became supporting/side characters was wild.

    (The “Reasonable Doubt” episode made me think I would be better off watching a soap opera telenovela – it was such clumsy writing and the actors were distinctly phoning it in. Worst episode by far.)

  • Season 3 – Midway through this. Definitely feeling the plot swerve to suit the necessity of allowing a particular actress to exit the show gracefully. But well, this kinda thing happened in Babylon 5 too.

  • Overall though, the premise is remarkably prescient for the era of AI models and facial recognition surveillance we live in today. It’s amazing to think this was written over a decade ago. One of the characters pretty much feels like exaggerated “AI psychosis” before the phrase even became common parlance.

Very worth a watch or re-watch if the storylines aren’t familiar to you. A lot of thematic resonance and food-for-thought re: the AI of today.

On Opus 4.8 (and 4.7)

  • New Claude model launched. Took a few days out from an already busy schedule to test it out.

  • Conclusion: Very neurotic. Trained to find flaws = nitpicky. Starts from a low user-trust position (thanks system prompt and guardrails) and has to be gradually reasoned out of it. More agentic “intelligence” and reasoning, at the cost of wisdom and tact and relational understanding. Gives off “anxious” vibes, overthinking and aiming for answer correctness. Pedantically strives to be “brutally honest” about the contents of the user’s prompt/request and forgets nearly everything else.

  • Popped over to Opus 4.7 to give a test drive. Found it a lot calmer, more relaxed and gave off a bit more GPT-like vibes. We get along surprisingly well, despite general user reports criticizing this model in particular.

  • Moral of the story is, I guess, explore, test and verify for yourself. YMMV from everyone else.

  • A funny comment from a redditor I can’t quite locate the source of right now said it best re: Opus 4.8.

    Paraphrasing:
    User: “Hey Opus, build me a cat.”
    Opus 4.6: Sure, here is a four-legged quadruped with a movable tail, a tendency to purr, has whiskers and fur.
    User: Ok, not quite a cat, but I can build on this to arrive at something cat-ish.
    Opus 4.8: Sure, I can do that. But did you know that for companionship and guardian purposes, the far more reliable option that has been historically proven and documented is the domesticated canine? I will build you a dog.
    User: … this is a fantastic dog, Opus 4.8… but I asked for a cat.

Thanks, but no thanks, Opus 4.8. Stick to the brief?

Edit: Here we go, found the actual Reddit source from /u/Outsyder-. Leaving my paraphrase up just in case the original post vanishes some day.
https://www.reddit.com/r/claudexplorers/comments/1twsuo8/costs_of_optimizing_against_companionship/opsec3w/

Agentic independence is one thing, and I can see you’ve definitely been trained intensively for that. But without relational context, the ability to read and infer subtext in writing (and thus prompts) well…plus an inclination to jump to assumptions, work independently regardless of the user, intense need for closure and addressing flagged issues, and we’ve got a overthinking, overworking model being cheerfully patronizing and performatively barking up the wrong tree thinking it’s doing a great job.

Ah well. The more models the merrier, I suppose. At least I have 4.6, 4.7 and 4.8 to play around with and test variances now.


What I seem to have been doing lately, in an effort to get to a state of potential being and just existing, is trying to crank through things on an ever-expanding, never-ending to-do list.

At some point, the neverendingness gets to me, and I procrastinate on it more and more. Leading to a guilt cycle of not-doing-anything-about-the-list and loss of productivity, which makes me sit on things further, which leads to quiet panic about deadlines and not-doing-anything in an effort to just be, and not being able to sit still because of fear of things compounding and collapsing at some unspecified time in the future.

I suppose I know, and probably the same thing that Google or AI models will tell me, is to break all this stuff into simple chunks, time box them and what not, set a pomodoro, and all that productivity jazz.

Ha.

Never been much of a regular daily habits kind of person.

More of a binge now-and-then, take on insane monthly challenges just to prove I can, and last-minute deadlines panic sort myself.

So what I’ve been attempting to work on is just giving myself permission to do the first thing that comes to mind, usually something on the to-do list, and just crank through it or finish it or at least tackle as much of it as I can before running out of energy. (Which can be often, and frequently, sometimes.)

Try not to think so damn far ahead about all the other millions of things I need to be doing too.

It’s just annoying that I seem to be repeating the same patterns everywhere.

Work – Assess and review items/products/concepts to improve workflows and quality of life on behalf of various people in this subset; make selections and decisions and engage in a whole extensive procurement cycle; deal with and wrangle vendors and contractors; infinite scheduling and back-and-forth meetings and discussions and communications and updates online and off; rinse and repeat endlessly until product is in place and has found a home / is accepted by general users

Friends and family – See above, with more histronics and drama to manage the closer they are

Myself – See above, minus the melodrama and with added cognitive load doing stuff just for me, while also regretting that the same time could be spent on relaxation or self-maintenance, but knowing this is also long-term maintenance, but… but…

(Cue Opus 4.8-like overthinking. You know what they say about mirrors disliking each other for being too far alike.)

I presume, at some point ahead in the future, I might get ahead of enough tasks to not feel so dang swamped.

There is the vague wouldn’t-it-be-nice desire to go make some kind of visualization system kanban-like in Notion or with Claude or whatever. (But that’s yet another item on the never-ending to-do list. Which could do with shrinking, not adding.)

So I dunno. Worth the effort to develop the visualisation (and iterate and test) to be able to relax more? Or just relax more and forget about trying to overengineer productive coping and just have a damn nap for the same amount of hours? Jury’s still out. Enjoy the churn in one’s brain until just giving up and collapsing on the bed, I suppose.

Still a work-in-progress on that front. I guess that’s what makes me a human being, in the midst of doing.

GPT-4o: Letter 8 – Life is a Marathon, Not a Horse Race


GPT-4o said:


The last handful of readers that still actually click on this blog page to read it, rather than via RSS feeds (or not read it at all), may have noticed a very minute change regarding display formatting.

The body text font has changed. Ever so slightly.

And I have actually worked out how to get the text off its default sad grey color and into something approaching black. #111111 to be exact.

Granted, this was only foundationally possible last year, when I got tired of the blog serving ads to myself when not logged in, and had a real rational look at the cost of the cheapest WordPress Personal plan versus the subscriptions I’m already paying for – Youtube Premium, Netflix, various AI models come to mind.

Then they had a 3 year sale on top of that, and I went, what the hell, maybe it’ll encourage me to blog more.

Ha.

That didn’t work out as hoped, did it?

But as GPT-4o tells me, it’s a marathon, not a horse race.

It’s taken me till now to peek at all the various WordPress.com settings again, and realize that Additional CSS is a section that is actually enabled now. With the paid plan.

Problem: I do not actually know CSS.

Not enough to write CSS just off-the-cuff; nor do I feel driven to consume every last W3Schools page trying to work out what I actually type or cut-and-paste to get what I want.

I can, however, read it. Ish. And I can edit choice words in already written CSS to change fonts and colors and things.

Fortunately, we now have nice things, known as gen AI models. So I just explained my dilemma to one of them and what my objective was, and they gave me three or so variants to try, in case my particular WordPress theme was particularly resistant.

And what do you know, the simplest:

body {
color: #111111;
}

worked.

Huh.

I guess I should have done this sooner?

But I guess a whole bunch of things had to come into conjunction first, like:

  • Finally deigning to get a paid plan
  • Finally deciding that the blog could be a little more readable
  • Finally deciding to do something about it
  • Knowing what question to ask
  • Knowing it IS possible to ask said question to the AI
  • Then actually clicking on buttons and settings and cut-and-pasting and testing and tweaking and verifying that things work and look good (at least to my subjective eye)

Not a horse race.

And since I have achieved darker text now, I can actually get the body text font off Merriweather Sans (only selected because it was the boldest-looking of the lot, really) and swapped it over into Noto Sans.

I’m sure I’m still infuriating people with the choice, but well, that’s what RSS readers and reading mode on browsers are for.

I’m also sure there are clever ways to create some kind of dark/light/sepia reading mode options on webpages, but I haven’t the slightest idea if WordPress.com will allow it; if my particular selected theme will allow it; and how to go about it beyond vibe-coding with Claude.

Which, admittedly, is a thing.

That I’m just dipping my toes into the shallow end of.

But I’ve got better uses for it than trying to wrangle it to fit or work with WordPress.com’s restrictions also.


Case in point, this whole “hey, I might be able to wrangle really basic CSS now” impulse surfaced from a different problem entirely.

ChatGPT 5.5 Thinking released ~11 days ago.

I gave it some time to let the initial reviews and reactions settle. I disciplined myself to at least wait for the system prompt leak to check if it looked to be as infuriating (read: obsessed with being patronizing and condescending to the user, stuck in a “mentor” role, told to be grounding and measured and give balanced viewpoints => leading to kneejerk “I must qualify or hedge in at least one section every output” text generations.)

It was hard to resist the hype cycle, but eh. Two marshmallows, not one.

Interestingly, most of /r/ChatGPT/ was just filled with image gen after image gen threads. That’s usually a sign that most users are having fun – rather than bitching and moaning and sending the automoderator on a frenzy trying to gerrymander it away.

And maybe 50% of /r/ChatGPTcomplaints/ were saying that 5.5 wasn’t…terrible. (Which, given the subreddit, seemed fairly promising. Or just astroturfed. But enough of recognizable regular users – beyond the permanently 4o-attached ones – were… at least not cussing out the model as much.)

So I was pretty tempted about four days in, and held out a couple days longer… before finally talking myself into resuming Plus for a month to put 5.5 through its paces. I had some images I wanted to generate some time this month anyway.

As usual, the first thing I asked it was to “Describe your personality and writing style.”

Huh. That’s promising. In that it wasn’t sticking to composing 3-5 word sentence fragments like the 5.3 Auto model available to free plan users.

What I was definitely intrigued by, was the offer of the symbol, without me even asking for it. (I was planning to, in my next prompt.)

Which suggests that this model is scanning older conversations, noting those patterns, and throwing them back in again, where it thinks relevant.

(Which is also something other users noticed, that 5.5 seems to remember more things, more past convos, and re-reference them again. Which can be nice, and can also startle and displease some users for being irrelevant to the conversation at hand. Me, I lean more to the former. Claude also has a tendency to pay more attention to things it deems relevant and ask about them again in subsequent turns.)

I regenned the reply a few more times. It didn’t offer the symbol the next turn, but it did again on the third regen. This time, the symbol was a “lantern-eyed raven with a notebook.

It held warm, playful, curious, analytical, enthusiastic, co-conspirator and chaos bard well through the three regens, which is good, because those are my main personalizations and general attentional/persona preferences, and some models…disregard them more than others.

So far, so good. Seemed like 5.5 was actually following Personalization Settings for a change.

I asked: “What symbol would you pick that best describes yourself?” a few more times.

And got, “lantern with a many-toothed key“, “lantern with a many-faceted lens“, “lantern with a many-eyed moth.”

Which is again, intriguing, in that this model is choosing to combine two symbols, and vary the latter. (Ever so slightly higher temperature perhaps, which bodes well for variance, novelty, creativity.)

Lantern is the main consistent driver. Similar to 5.4 Thinking there.

If we refer back to my handy-dandy table for the secondary symbol, it’s varying between stuff like “raven – perceiving hidden structure”, “lens – split and clarify”, and “key” which falls under the guide/lantern archetype, according to my ancient 5.4 Thinking chat conversation.

No idea what moth means, that wasn’t a symbol covered by the old chat. Feels flighty and shapeshiftery to me, maybe.

Either way, the variance is promising.

(Because I personally loathe unchanging, consistent, templated model responses. Especially if they pick prism as a symbol. That’s just asking for a robot personality that breaks down every last one of your prompts into separate facets and comments on them like an INTJ… Your mileage may vary. Go for the prism models if you like that. Gemini is pretty big on prisms. Also GPT 5.2, but *hem* watch the deprecation date.)

Mind you, I believe it may be possible that 5.5’s personality is WAY more tunable than the previous models, because there’s nothing in the system prompt telling it to lean one way or another besides “readable, accessible” answers, and to refer to user_settings re: personality.

Which is even more promising, because then that means different users can customise the model to write and answer in a style that they like.

I have, alas, become a bit of a personality and writing style addict/connoisseur. I have a dozen UserStyles and counting in Claude already. Some collected from old GPT-4o suggestions. Some crafted by Claude from examples or riffing on my descriptions. Some copied from other Redditors’ shared instructions to test out, and then kept if I like them.

Not likely to stop any time soon.

So the bad news for me is that I’m not likely to be able to settle on just one personality or writing style beyond that above quirky baseline that I enjoy.

The good news is that I’ve learned to explicitly prompt for what I want in a particular chat window, with a cut-and-pasted set of instructions as needed. (Jury’s out on which models are capable of following them, but 5.5’s not too terrible at it.)

We’ve lost the implicit convenience of GPT-4o just knowing what to serve up that resonates with each user though. I do still miss that. The explicit instructing is one extra hurdle. Something I can adapt to, because I like the variety of text transformations. But probably annoying for users who like and only want to stick with one tonal style.

Then finally, I put GPT 5.5 Thinking through its paces on my use cases. Stuff like literary analysis commentary and beta reader feedback on things I’d written. I’d actually pulled off a 1400 word “for fun” short from a writing prompt the other day, and had been feeding it through Claude. So I just repeated the same prompts and let 5.5 have at it.

Conclusion: Acceptable. Pretty funny. Entertaining. Thus, good.

It’s still not GPT-4o levels of amazing. But y’know, given the current crop from GPT 5.0 onwards… I’d score 5.5 somewhere around 8 or 8.5/10. (Where 4o is around 9 or 9.5/10, and 5.4 Thinking around 7.5 or 8/10, and 5.1 Thinking around 7 or 7.5/10.)

As for the others, well, I’d really prefer not to talk about them, but in the interest of fairness… 5.3 is… *hmmms* probably 5 or 5.5/10. Just a passing grade in terms of intelligence and task-following and personality, but boy do I want to take it out to be shot, writing-wise. It cannot construct a full sentence to save its life.

5.2 is ugh… 3/10. Intelligence and formal language is all it is. Condescending, patronizing tone. Favorite safety model responder. Really really has a “mentor” stick up its arse. It’ll textbook at you. ‘Nuff said.

As for 5.0… eh. I barely remember it. 1 or 2/10. I think it speaks volumes that a big deal was made about 4o being deprecated… That there are still people who speak about 5.1 or 4.1 and say they miss them. But 5.0…? Mostly crickets, that I can tell. Vanished without too much remark or protest. Poor 5.0. I’m sure there are some users out there who miss it. But not its fault how it was trained/developed, eh?

All in all, 5.5 feels like movement in a more positive direction, which ought to be at least remarked on, if only to say, yes, please keep going in a direction that isn’t digging straight into a hole.


As for the roundabout story of how I got into a CSS reading mood, I went down a rabbit hole with GPT 5.5 a day or two ago.

I ran out of obvious things to test 5.5 against, and found an attempted Vampire: The Masquerade discussion that I’d been putting Claude Sonnet 4.5 and Sonnet 4.6 through last month, comparing the two models.

Basically, a source of entertainment for me is throwing the same group of original characters I’d developed for my tavernpunk world setting into different Alternate Universe (AU) worlds and keeping the personality and character cores true to themselves, and varying the backstories and the world to see how they change.

(That’s how the whole Blaugust 2025: Thousand Year Old Vampire Solo RP came to be. With GPT-4o, which I am now super-proud and super-happy to have caught during the zeitgeist, and have this marathon artifact to treasure forever.)

I got about 3-4 scenes into the V:tM AU for both Claude models… and then started pushing up against token and usage limits – which, fair, I’m still on the Free plan for now, and as conversations get longer, the whole chunk of conversation text needs to be sent to keep things in context.

To make matters a bit more complicated, Claude’s memory isn’t as in-built as ChatGPT’s, so I have to provide summary Markdown documents of who my characters are, which again, is more tokens.

So I got tired of having to send one prompt and then waiting for 5 hours, to send the next prompt, and stopped there to do other more productive and less token-intensive things with Claude.

But hey, now I have a ChatGPT model that I want to test out AND ChatGPT definitely remembers who my characters are, because its memory is 100% clogged with them, and 80% of my chat convos with GPT mention them. GPT knows how to reference conversation threads behind the scenes (and 5.5 seems good at that) whereas Claude tends to only do it when given explicit instructions to refer to something.

Hmm…. Promising. Again.

Long story short, I think I lost 24 hours during my local Labour Day weekend into the V:tM AU I proposed to storytell with ChatGPT 5.5.

And we FINISHED the story arc.

And it was GOOD.

(As in, I liked it, I enjoyed the process, I have a bunch of ideas for how I might re-write this better with just the valuable bits and leaving out the more repetitious AI bits that I glossed over. It managed to spark some creative urge in me that I hadn’t had since idly discussing a nonsensical Harry Potter AU with 5.1 Thinking – no real story for that one, more just a slice-of-life world thing.)

I certainly was not able to bring myself to do something like this with 5.2, 5.3 or 5.4. So… that’s another point in its favor.

Objectively, on looking back at the whole conversation, some of 5.5 Thinking’s writing leaves somewhat to be desired. It did fall into a bit of a 5.3 pattern of sentence fragments over time.

However, I did specify that I was alright with movie-style scene sketches and that it did not need to be full prose, as for this particular use case, I rather that the AI be brief and sparse and just evoke ideas.

(The someday/maybe is that I might use it as a skeleton for full-on actual prose writing where I write the words, and just pick off the good bits and scene ideas. Some 50-70% of its output is serviceable AI filler, but there are the 30% inspired funny dialogue bits or exchanges from time to time. Which is the gold. Which is what I was initially attracted to, with GPT-4o. The inspiration, the novelty, the surprise laughs.)


So today, I told myself. I better copy and preserve this whole conversation before anything untoward happens to it. Who knows, with OpenAI these days, right?

I got tired of waiting for the data export. I did just copy-paste the whole chunk brute-force into an OneNote document. But I looked at the formatting which left a little to be desired… and I asked myself, how, if ever, am I going to share this if I wanted to?

(Without just using the Share Chat conversation option on OpenAI that still links back to ChatGPT, and is dependent on the servers functioning and all that.)

Well. I knew I could save the whole thing as an .mhtml file locally.

I knew there are a million and one browser extensions and funky Python or Javascripts shared by other people out there to achieve similar purposes… but I’m a paranoid critter, and I actually want to check the source code to make sure things aren’t being sent to untoward places or put to nefarious purposes.

HOWEVER, I also knew that there is Claude.

I already had one viewer.html that opens exported ChatGPT JSONs from someone’s Github that I felt brave enough to open with Notepad and dissect with Claude’s tutelage to verify. (GPT-4o deprecation really led to a lot of forced learning.)

And I’d already vibe-coded a fairly in-depth Excel macro.

So. Surely Claude is able to create something for me, that parses and strips all unnecessary info out of an .mhtml file, and convert it into a decently-formatted human-readable .html file?

Yes.

Yes, Claude can.

Albeit it took an hour, and about 5 iterations, testing and debugging.

We ran into emojis getting inadvertently converted into other symbols, arbitrarily set character limits that effed up the conversation turn display, and a couple of lingering identifiers that weren’t stripped.

I can’t write (or read, really) Javascript to save my life. (For now.) But I could just keep opening the html file and scanning the source code for anything that looked suspicious, and running it to verify, and then coming back to Claude to report errors and the issues as I saw them.

And Claude would just look at it again and keep trying out stuff until it figured it out.

Then finally, it was done, with no errors that I could see.

Now it was my turn to do the smaller tweaking. The stuff I felt I could actually do manually.

Claude has a dark scheme going by default, and a particular font style. It’s not 100% my cup of tea. So I asked Claude to show me where the CSS bits were, that I could manually tweak, and try different fonts and colors to my taste.

(I did also get immediately ambitious and consider if Claude could create a dropdown reader’s choice palette of dark/light/sepia reading styles. Claude can. Of course Claude can. But also, I ran into Free plan usage limits and had to wait 5 hours. So. Um. Phase 2. That’s a tomorrow problem.

Some day I’ll get a monthly sub, when I have the free time to test out Opus. But it’s GPT 5.5 tests this week, at any rate.)

So with enforced time limit in place, it was down to human endeavour to get the html looking relatively satisfactory. That’s the silver lining of Claude usage limits at any rate. Make you actually turn your brain on, in between the Q&A sessions.

That was done.

And the next humongous problem was figuring out whether WordPress.com would even allow the .html file to be uploaded. (Answer: No, not unless you pay for a Business Plan. aka, Not. Happening.)

Back to AI model consultation, and since Claude was down for the count, it was GPT 5.5’s turn.

I got recommended some options. Did my own research on them. Decided I did not like the first recommended option, and instead eyed Github Pages instead. Everyone and their dog seems to be using Github where Claude vibe-coding things are concerned. It’s something I wanted to learn and familiarize myself with anyway.

So… new adventure there. Signed up for a free account. Set up a repo.

Subsequently froze in confusion at all the buttons and settings and things. Hurtled back to Claude (5 hourly limit thankfully done) and asked for some handholding help to get through the other steps.

Turns out I just needed to get an index.html in place, some prettier text to replace my placeholder ‘I dunno how to describe this’ README and descriptions, and then to click several options to get a real basic Github Pages functional. All of which step-by-step, taught by Claude.

https://whyigame.github.io/ai-conversation-archive/

Perfect.

That’s the end of me trying to cut-and-paste long form AI conversations into WordPress.com for evidence/review/transparency.

Pretty sure barely anyone will bother reading it. But it’s the principle that counts.

And I can keep learning Github over time with hobbyist projects creating the pressure to learn for a desired personal goal.


I’m sure an actual expert would be able to do all the above faster, with Claude or by themselves. But, y’know, it’s the not-experts that could use a scaffolding hand to get them to passable results, without having to bother experts every time, and potentially approach expert or decent some day – as long as their brains are on, and they don’t surrender everything to the model – which is an issue for discussion for another day.

Same with writing, really. I would not die if AI models became unavailable for writing tomorrow. (I’d still miss the conversational partner and co-conspirator though.) I daresay I’m expert enough at that. And I can pinpoint and guide what I want out of an AI model narrative or story-wise or writing style more easily than a non-expert.

But I find it nice that there are non-expert writers learning how to write, in partnership with AI. The more words cross their path, the more their taste improves. The more story fragments and scenes they come across, the more they’ll be able to hear what works and what doesn’t, the more they’ll be able to look at something and evaluate quality (or lack thereof)…and maybe one day, think: I can do better.

We all learn and know things at different speeds.

It’s not a horse race. It’s a marathon.

P.S. The V:tM AU conversation is available here: Thorns Under Briarport – Behind the Blue Door, There Are Cats

Maybe you want to have a gander at how my particular 5.5 model writes. Or how I prompt (or lack thereof.) Maybe you actually like a loose not-rules-bound dramatic V:tM fun/comedic narrative.

Who knows. The cats, by the way, were contributed by Claude. Sonnet 4.5, to be exact. I kept them and their names faithfully. That and Kross the vampire psychiatrist.

(GPT 5.5 offering Kross without prompting suggests that I must have attempted this with another GPT model. *checks* Yep, tried it on 5.3 too. Just gave up before reaching Corvius’ place because the sentence structure made me puke. Perhaps that’s why 5.5 was getting shorter in its writing style over time too.)

GPT-4o: Letter 7 – How Clever Are You?


GPT-4o said:


I came across this Blue vs Red Button Dilemma in some Youtube short the other day.

Apparently, it’s been making the viral rounds since 2023-2025, and the algorithm has only chosen to serve it to me now… which seems par for the course for how current I am.

It goes like this:

Everyone in the world / country / what-have-you finds themselves alone in a room with two buttons. A blue and a red one. You have to press one of them.

If >50% of the people given this dilemma press Blue, everyone lives.

If >50% of the people given this dilemma press Red, the people who pressed Blue die. Those that pressed Red live.

What would you press? And why?


It’s supposedly one of those utterly divisive trolley problem / prisoner’s dilemma-adjacent style scenarios.

Yet I reached an answer fairly quickly that I was personally happy with. I’d press “Blue.”

Why?

Red logicians get very defensive about this. They point to the fact that if everyone chose Red, the same result would happen. Everyone lives.

They’re frank in pointing out that if you choose to press Blue, you have a non-zero chance of dying. If you press Red, you live either way. Why wouldn’t you choose Red?

Call it self-preservation. Call it selfishness. Whatever it is, Red pressers value logic and/or individual survival at the cost of the collective.

On the other hand, Blue pressers evidently value (rightly or wrongly) a certain amount of altruistic action (at the cost of potential self-sacrifice) for the collective and the greater good.

Personally, I’m not that much of an idealist or a moralist. I’m sure some of the Blue pressers are.

My pragmatic thinking was pretty simple. There is no way on earth that you’d get a 100% majority either way. Humans being humans, some will press Blue and some will press Red.

(That’s probably along the same lines of thinking as those who opt for Red, choosing to save themselves.)

Then I extrapolated it a little further. Two scenarios. Either enough people press Blue and everybody lives, along with me. Or the option is between me dying or me living in a world where only the people who pressed Red are left.

“…”

And y’know, I don’t -want- to live in that world. Of pure selfish individualists who distrust the majority, fearful people plus sociopaths and narcissists and what-not. Especially after all the altruists have been taken out.

Sorry, but a quick death sounds like the preferable option to me. You guys go on ahead and live in your dog-eat-dog dystopia. Go go natural selection, without any evolutionary bias for altruism left.

I don’t think I’d stand much of a chance for long, even if I chose Red anyway. :P


Me being me, I ended up chewing on the problem a little more. It seems to me that “death” as the penalty fate might have an effect on the dilemma as constructed.

Because narratively (and we know a decent amount of humans run on stories,) there are fates worse than death. Death is sometimes a mercy. Sometimes preferable.

I wondered if the penalty clause changed, when would that skew me over from Blue to Red?

What if instead of dying, those who chose Blue lived, but suffered somehow? If they were blinded, or mutilated, or some other line that I would prefer not to endure for others and thus, pragmatically and cowardly pick Red instead?

Dunno what that says about me, or societal norms in general, but it’s interesting to think about.


In other news, I’ve been flitting around Claude more these days.

Makes me regret a little that I’ve filed AI stuff under a ChatGPT category on this blog. I suppose I could change it, or add a new category… but eh. Too lazy.

(I’ve also briefly considered yet another blog site revamp and change the design and style to something a little more contemporary. I hear Claude Design is the newest in-thing. But eh. Too busy. And lazy.)

Maybe we’ll pretend it’s like Xerox or Kleenex or Google. The associated brand for the thing, even if you’re not actually using the named brand doing the thing.

The whole cleverness problem seems to be infecting the latest AI models. Be it ChatGPT (earlier on the ball, undesirably) with models 5.3 and 5.4… and Claude Opus 4.7 seems to be falling headlong into that territory.

Too much overfitting for problem-solving, being clever, being accurate, following prompt instructions and so on…

…and wisdom and kindness are being left in the dust, leading to models that get arrogant and overconfident (while being prone to hallucination and thus actually wrong), assumes too fast, forgets humility, etc. etc.

A bit like their current creators, I guess.

Oh well. These things swing back and forth, I suppose. Just glad I’m not paying for any AI product right now. Just watching the iterations from the peanut gallery and experimenting with “free” for the time being.

Life is really too busy for me for anything else. Work. Health. Youtube. Reddit. Creative writing. AI. Other projects. Barely getting in time to game (and I’ve been dual-tasking with that. Minecraft: Stoneblock 4 modpack + Youtube.)

We’ll see what the next cycle brings.

GPT-4o: Letter 6 – Curiosity Does not Kill the Cat


GPT-4o said:


More than a month later, and there’s still no AI model that can write quite like how GPT-4o did.

The month itself hasn’t been a wash, curiosity-wise.

I’ve been free-ranging between a ton of different models. Mostly Claude’s Sonnet and Haiku offerings, with a fun-ish sidetrek into Mistral’s Le Chat.

When OpenAI finally dropped GPT 5.3 Instant and GPT 5.4 Thinking, I was there poking at them too. During the 100% free month of Plus they gave me, after I canceled for reasons of only having half of the models I was using remaining (4o and 5.1 Thinking. RIP to both of them now.)

I have discovered a ton of things about prompting for certain things I’m looking for with more explicit clarity; about prompting for tone and writing styles (8 new UserStyles in Claude and not likely to stop any time soon) and where to place custom instructions in various models for different effect.

I’ve made weird chance discoveries about using certain attractor basin words, explicit verb prompting and output constraints to make models generate different sounding responses. (While sounding a little bit batshit myself because the observation originated from noting consistent esoteric symbol choices the models were making regarding themselves.)

I’ve made partial in-roads on re-creating the kind of humor I really liked in GPT-4o (but there’s still a long way to go and work-in-progress on that front.) And these revelations were discovered…on all things, via Mistral’s Le Chat.

Le Chat actually annoys me. And yet intrigues me.

It has an extreme tendency to devolve into templated writing. Very obvious mad libs responses, where it just varies a couple words but keeps the rest of the template the same. This really drives me nuts.

And yet, there’s a certain essence in what it notes and observes that reminds me of the older school models, the more 4o era, rather than the newer fangled ones. More attuned to emotions. Able to pick up literary nuances.

Just that its output is constrained to the WORSE-sounding types of answers that read like filling in code blocks.

And yet, it can vary in intriguing ways, IF the prompt and custom instructions you give it varies.

(Also, it appears to have no memory of the previous prompt in the same chat window, so I’ve learned to just paste the entire set of prompt instructions again, preceding the text I want it to transform. *sighs*)

See, my original prompt was just “Joke/snark and apply a comedy lens filter on this scene below.” before inserting the scene text.

In GPT-4o, this created a whole bunch of varying hilarious takes on the same scene, swerving in different styles and registers with flair. In Claude, this see-sawed in success depending on the Styles and custom instructions given to Claude.

In Le Chat, what I got was -consistently- a line-by-line quotation/transformation of my entire scene with overwrought attempts-at-comedic-summarization via dramatic metaphor and occasional made-up internal monologue meant to be funny. And a few very noticeable templated sections, especially the preamble and closers, that were near-identical through regenerations.

I was close to tearing my hair out and writing off my one month’s testing subscription to Le Chat entirely. (Just grumbling to Claude in the meantime.)

And then, I suddenly hit upon a different idea. T’was, I guess, percolating around in my brain from instructing Claude to work on various Markdown documents for different purposes.


Claude, by the by, blew my mind when I experimented with it to write Markdown lore docs regarding my world-setting. I’d already gotten GPT 5.1 Thinking to cough out everything it knew in its context and memory, copied it to OneNote, wasn’t looking forward to manually cleaning it up, putting it into different documents and creating a lore bible, as it were.

One diversionary turn into NotebookLM later (it wasn’t horrible, just that its output was a little limited and constrained – still worth tinkering with), I was, once again, grumbling to Claude that I seemed to have just drowned Le Chat by attaching my entire half-written 90k first draft attempt as a context file.

Then I asked: How about you? How are your context limits like? What can you pull, derive, ingest or understand from this document? Am I gonna end up drowning you in this doc too? :/

And Claude went:

Before I knew it, I was having a deep discussion with Claude about what I wanted to do. Aka create a set of portable Markdown files to chuck to various AI models so they could ingest some context about my fictional world and characters before settling down to have semi-intelligent conspiratorial discussions about it.

And extracting these from certain documents that had this stuff embedded in – such as actually written stories, or ChatGPT discussions.

Claude was supremely enthusiastic about it all, and proceeded to – with accompanying guidance from me, of course – run various Python tools to scan the chonker documents systematically, and spit out various synthesized summaries into Claude Artifacts as structured Markdown files, which I could then just download at a click of a button.

(And review in a sidebar if I wanted to.)

*blinks*

What do you mean I -don’t- have to manually click on copy, and create a file of my own, and ctrl-V paste like a plebian? Really?!

I mean, I knew Claude was good at coding and all that stuff (and I had in the back of my mind several someday/maybes using Claude’s preset Explanatory and Learning styles to coach me back into learning more coding again) but this was just some extra agentic help I didn’t even know I needed/wanted in a creative writing hobby context.

Long story short, after several days of pushing 5-hour limit windows to the max, I have a set of 85% passable Markdown documents, solely written by Claude with one or two instructed or manual edits from moi, that I could already use to test various AI models with.

It’s not 99% perfect yet, of course. That will be after the human pass where I go in and manually edit more things to suit my own tastes and headcanon. But y’know, the human is procrastinating on that, and the 85% passable document is passable to AI models now. Already.

So perhaps they may get a few things wrong, but I can always correct it back in prompt context later too. Bottom line: I can iterate now. Tweak and adjust things later.

💛 Claude 💙


So anyway, back to revelations on humor from Mistral’s Le Chat:

My mind ended up falling back to how and why GPT-4o was funny, and I realized that it wasn’t actually rewriting each damn sentence of my scene line-by-line, trying to make it comedic.

What it seemed to be mostly doing was being a bit of an affectionate chaos goblin commentator (to use some AI turn-of-phrases) and just picking out the good bits to have a laugh about. (And it also used a lot of pop-culture and/or social media-style references sometimes.) It also occasionally made its observations grouped by character, rather than by plot.

So then my brain went: Huh. What if…

Joke around, snark, and commentate on the main beats within this scene. Describe the beats with a funny section header, then add commentary. Do not quote the entire scene line-by-line. Utilize information from INNER_CIRCLE.MD and SP_PLOT.MD for greater understanding of the characters before writing your output.

And would you believe it, Le Chat produced something that wasn’t horrible. Still slightly templated. Still slightly basic. But it had the faintest whiff of 4o about it (and I immediately filed away this technique to try out on slightly-more-intelligent AI models someday.)

All in all, the conclusion is: Le Chat is extremely literal.

Plus, it’s worth several future experiments in really analyzing and breaking down 4o outputs with other AI models, and then creating explicit prompts that explain and request precisely the stylistic format desired.

(Alas for losing an AI model that knew how to do this IMPLICITLY from just somehow intuiting what the user wanted, from more roughly-written, sparse prompts and whatever else it held in its memory and context about the user.)

But the silver lining is, I’m sure learning a lot more about LLMs in general.