I am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.
+1. First thing I look for in a model announcement thread. I actually came across this one an hour ago and was sad there were no pelicans yet.
It's a decent heuristic because the better models generate better pelicans. That's all. Nobody sane is going to make a bet on a model based on a pelican. But it's cool, it's tradition by now, and it's a semblance of a good first impression for new models.
Thank you for posting them (as well as all your other AI takes). It's appreciated here!
I feel like the human brain massively overweights negative feedback over positive, and that's even after accounting for the fact that internet discourse tends to mostly surface negative comments (whereas the enjoyers stay silent). By default I have to try hard not to take it personally whenever someone leaves a negative comment about my work.
So just doing my bit to say I appreciate your commentary on so much of the fast-evolving AI landscape. Helps me orient :)
If enough people agreed with you, simonw's top level comment would be grey. It isn't. Not every comment is for every person, that is fine and normal. He didnt hijack a popular thread in here to make his post. He doesnt have a tin can in your face shouting from a soapbox that makes it hard to ignore. Flag or ignore are reasonable options that can be used.
People aren't supposed to upvote or downvote posts for these kinds of reasons here. Most people are downvoting far too much on this website, and it leads to significant echo-chamber dynamics that are worse than even reddit. The pelican test continuing to be taken seriously is a great example of that kind of echo chamber.
"He doesn't have a tin can in your face shouting from a soapbox that makes it hard to ignore."
He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated. Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts. Look at how bloated this very thread is right now!
Also, a lot of people unironically are whining about him because of sour grapes. Pay them what Simon is likely making, give them as much mindshare/attention as Simon gets, and they wouldn't be so mad.
Anti-incumbency bias and anti-elitist attitudes are good actually.
> People aren't supposed to upvote or downvote posts for these kinds of reasons here.
Agreed, but scrolling past or hitting - were apparently off the table for the complainers. So flagging was yet another tool in their toolbelt and I bet a powerful one at that. If simonw's pelican posts routinely went dead from flagging, he would not make them. You know that, I know that.
> The pelican test continuing to be taken seriously is a great example of that kind of echo chamber.
You can try and support that argument if you like. But I would implore you to realize that it has been had many times recently and the other side does in fact find value and do not see it that way.
> He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated.
Users upvote the pelicans because they find it interesting. they arent paid trolls or simonw fanatics.
> Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts.
They can learn... it's called hackernews. For those interested, that is what the [-] link is for above the comment. Use it and move on.
But, also, as said, you're an industry (and foss!) veteran, so I find it impossible to believe that you haven't had your fair share of baseless bullshit being thrown at you, and with that, you gaining a persona that will not be hit by that, because it clearly knows that it is in fact bullshit.
Unless of course it doesn't really know that with certainty.
As said, I would _love_ to give you the benefit of the doubt, because you might just have a stressful day or whatever, but content marketing is literally your whole thing by now.
It is impossible for me to do that with a clean conscience.
Your blog front page currently opens with
> Earlier this month I hosted a fireside chat session at the AI Engineer World’s Fair with Cat Wu and Thariq Shihipar from Anthropic’s Claude Code team.
That is not what "some rando foss maintainer we are morally obligated to be soft with" does.
I'm a professional blogger now. I still also work on open source software. I'm even fine being called an "influencer" (shudder), but I take offense to accusations of unethical behavior.
I think very hard about the ethics of what I'm doing and how I can best use my "platform" (shudder again) in as constructive a way as possible.
What do you want from him, seriously? Anyone who follows the scene knows that blogging about AI (and participating in conferences etc.) is what Simon does nowadays.
There's nothing shady here. The disclosure is front and center on his About page on his website.
He's not spamming you. It's one short link, sometimes a link to a first-impressions post. It's interesting and useful for me and the other commenters who keep upvoting his comments. Why are you so antagonistic?
At this point, it's kind of a hackernews thing. Simon posts them as a single comment in the relevant thread. It's okay for this place to have a little bit of a sense of community, and you can just ignore the comment.
But how else am I supposed to know when we've reached AGI, until I see an absolutely flawless pelican?
All of the pelicans so far have had really weird flaws / quirks so I am always a little interested to see how well these models perform at this task, since I've seen all the past pelicans and have some anchoring.
Seeing a truly flawless pelican would tell me that the model has true visual reasoning capabilities as well as good taste.
I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it.
It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
Sorry mate, but you sound jealous in all these replies that the Pelican domain isn't your gig. The below is as labored as nitpicks ever get:
> It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
I thought the Gemini 3.5 Flash Lite response was quite telling myself. I personally like the Pelican SVG test, to me it is still a charming snapshot of model performance anecdata. No one would argue it's rigorous but I don't think it was ever intended to be.
I get people burning out on the pelican SVG test alongside the rest of the AI burnout, but I guess for myself I'm just choosing to keep enjoying it while I still can.
Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.
Here is a different opinion. I’m always looking forward to see the pelican whenever a new model is released. It’s plain simple to understand, memorable, subtle enough in terms of details, and my favorite part is that you have been doing them consistently long enough for it to be useful for comparing almost everything with anything.
Disagree. They're a nice tradition, but besides that, they're a useful way of eyeballing improvements. I realise labs are likely to be training for Pelicans - but if they're all training for them, the differences in the results are as indicative as they were before labs trained for them.
The 3.6 Flash pelican is just about the best I've seen.
Yeah. It's something I can do myself in a couple seconds if I want, also on more varied SVG scenes. If this is going to be a benchmark people turn to I'd like to see more effort put into it than just a one-sentence prompt.
Your comment reads very pedantic with a hint of jealousy. The pelican and xbox controllers are great ways to see how well it can follow direction dealing with svg a difficult format for LLMs to use and testing their spatial vision awareness.
I'll split the difference. When it's a blog post there's usually an interesting observation or two, but if it's totally automated? Maybe just do the ones with a post.
Honestly, i've received a formal MSc education in the hardware aspects, including for designing embedded electronics products. Spent the most part ofy career in the software industry designing enterprise software and feel like i never needed to use them, except maybe early in my career when i was reviewing tech stacks and determined that .NET would be among the winning horses, precisely because it'd take care of that for me almost all the time.
What i see today is the opposite of what you see : product owners not knowing a thing about software engineering but being able to vibe code prototypes handed over to the dev team are rock stars.
They are closely followed by senior software developers having more of an architecture & design background than a low-level computer science background. Most businesses are looking for builders these days.
Where what you say may converge with my observation is that to be able to do to things such as proper database query optimization, even using AI assistance, you need to be able to understand the concepts of working memory set, cache misses etc...
I've found huge problems, like database servers being grossly underprovisioned (like, 60% cache hit, 4gb RAM server for a 700gb dataset with an 50gb circa hot data set). SSD were used and only latency was measured, so no one realized how problematic the situation was (including a consulting shop they hired to help them manage their DBs - backup, maintenance etc...).
However, having a high affinity with hardware is not a driver / computer science of hiring decisions from what i can see in the enterprise software world. But it would make sense for it to become the case within 10 years. I suspect that you work in a niche where performance optimization matters a lot.
> However, having a high affinity with hardware is not a driver / computer science of hiring decisions from what i can see in the enterprise software world
I think the way I worded it was maybe a little too close to just being about hardware, because performance do matter a lot in the energy industry. I do think it applies to SWE in general. You mention .NET and I've met C# developers with years of experience who couldn't tell you the difference between IEnumerable and IQueryable. I've met even more experienced Python developers who don't know what a generator is. Stuff like that, not having knowledge of the tools they use. I guess you could argue that those are bad developers, but I don't personally think that has been the case for most of them. Still, you'd rather have someone who thinks about these things rather than eventually using batches once they run into memory issues.
I also think these changes are appearing faster in non SWE enterprise. As you said, product owners who are AI explorative (for the lack of a better word) are rock stars. We see this a lot in our finance and risk departments, where domain experts now write fairly decent software with AI. My team has build them tools so that they build things the same way, use the same developer setups and pull the pre-approved external packages or are offered alternative ways of doing things. A few years ago this would've been done by these domain experts "ordering" the software they needed from our SWE team, and if they hadn't already been mostly laid off due to Putin's invasion of Ukraine changing the markets, I belive they would've been now because of AI.
Because frankly, a lot of the software that gets produced in these areas, don't need computer science, until it does, and the domain experts can make the software they need so much faster than before by vibe coding it. From my perspective it's not that much of a difference in the quality of the code that gets produced. I also had to help with performance and security when we had more software engineers on staff. Though now I do it more through writing and distributing AI agent applications rather than writing a C binary or optimising the code directly.
That's true, although, if you look at them, you wouldn't notice. The only mention of JVM you can see in the IDEs is in the About dialog, and the IDEs install and run their own OpenJDK, so no JVM has to be installed globally. Almost as if they were a bit self-conscious about using such an "unsexy" architecture...
You mean a JRE, because the whole JDK contains a bunch of things you're never going to need.
Mind you, a default JRE redistribution makes your app at least 100+MB. Using jdeps to strip out unneeded things is a good idea if you want it to get down to 25 ish MBs.
Yeah, saying JRE is a bit of a shortcut since you're supposed to jlink & jpackage & jdance &jpray to get a slimmed down JVM released with your app, but it's closer to what would be a JRE than a full JDK
> the ability to quickly and safely change significant parts of the code and product.
Hum, this reminds me something... "O: open to extension, closed to modifications". Old things are new again.
From context efficiency with approaches such ad DDD and clean architectures, all the way to items such as this one, AI is not creating new tradeoff, it just acts as a multiplier, multiplying productivity for teams doing things right, and multiplying debt for teams having a low quality bar as far as design and architecture are concerned.
Exactly, it's funny how most Americans have no self-awareness on this topic.
And beyond VCs, which are like massive subsidies funded by printed dollars to which no other country has access, even in industries like electric vehicles, Chinese total direct subsidies to their EV companies are like $5bn per year, while the the ones provided by the US to their auto manufacturers are in the range of $50bn per year.
I don't think the US are cheaters or are doing something bad. But i do think that this propaganda about China flooding the market through "overcapacity" and subsidies is very dishonest and needs to stop.
The CCP's recent yearly subsidies for China's EV industry ranges from $30 billion - $46 billion per year compared to $10 billion - $20 billion per year for the US. A conservative estimate of total subsidies lands China at about $230 billion while the US was at $70 billion.
That's a big ask. This kind of harness usually contains plenty of proprietary insights about their business. And also, nowadays, a good harness is a major competitive advantage.
Your hostile tone is unfortunate, especially since my post was actually friendly. I was just trying to point why it is very likely the OP won't give you what you're asking so you're not left confused if he ends up ghosting you.
Many people use the term harness to refer to the agent coding software (eg. Opencode, Claude Code...), i use this term more broadly to refer to the environment (set of skills, system prompts, constraints, memory, hooks etc...). What the OP is referring to is not just one giant skill. It's usually a comprehensive ecosystem of skills, bespoke tools to make certain agent tasks deterministic (eg localization), and so on.
I've seen someone post Github repos in this thread, these can be very useful especially if you use the same tech stack, but you won't reach the level of productivity reported by successful teams unless you invest substantial time to build your own harness. But the way to do so is to do it progressively : start with something simple to address the need you have on day 1 . And then, turn recurring prompts into skills, turn recurring coding patterns and coding style recommendations into guidelines, turn repetivive tasks for which the LLM tends to build a python script that it occasionally gets wrong into a deterministic tool documented in a skill etc...
And after a couple of days, weeks, and months, you'll have a very dependable harness giving you optimal productivity, without needing to invest weeks of work upfront or take the fun out of agent-assisted coding.
> I wish these breathless blog posts would actually try to be more didactic.
Especially from a company actually trying to promote AI use, Mr. Occam says hiding such details is best explained by them not actually being that impressive.
As someone who uses quite a couple of different AI providers (codex, glm, deepseek, claude premium among others), i've noticed that claude tends to move too fast and execute commands without asking for permission.
For example, if i ask a question regarding an implementation decision while it is implementing a plan, it answers (or not) and immediately proceeds to make changes it assumes i want. Other models switch to chat mode, or ask for the best course of action.
Once this is said, i am not blaming Anthropic
For that one, because IMHO the OP has taken a lot of risks and failed to design a proper backup and recovery strategy. I wish them to recover from this though, this must be a very stressful situation for them.
All the models I have used will frequently jump ahead a ton of steps and not verify any of its assumptions. From generating a ton of code output I didn't ask for, to making a ton of assumptions about what I'm working on without appropriate context.
Yeah, /plan is the only way I can work with them now. Too much "helpful" crap I didn't ask for. Having nightmares of former coworkers who would want to refactor 80% of the code base for a 3 line change. AI doesn't subscribe to "if it ain't broke, don't fix it."
I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend.
I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long horizon tasks. One thing i've noticed is that it is particularly good at following instructions and guidelines, even deep into the execution of a plan.
To me the only problem is that z.ai have had trouble with inference : the performance of their API has been pretty poor at times. It looks like this is an hardware issue related to the Huawei chips they use rather than an issue with the model itself. The situation has been substantially improving over the past few weeks.
GLM5.1, GLM5-Turbo and GLM5v are at this point better than Opus, Codex, Gemini and other claude source models. We have reached a major turning point. To me, the only closed source model still in the game is codex as it is much faster at executing simple tasks and implementing already created plans.
Try GLM5v for your PDF work, it's their last generation vision model that has been released a couple of days ago.
Does anyone have inside info on what these Huawai chips look like? I know Google has a Torus architecture unlike Nvidias fully connected one. Maybe it’s a similar architectural decision on the huawai chips that leads to bottlenecks in serving?
>For AI computing, the Atlas 950 SuperPoD, powered by UnifiedBus, integrates 64 NPUs per cabinet and can scale up to 8,192 NPUs, delivering superior performance for large-scale AI training and high-concurrency inference.
Plenty of other providers that offer much faster inference on GLM-5.1. Friendli, GMICloud, Venice, Fireworks, etc. And can be deployed through Bedrock already as well. Will probably be available generally in Bedrock soon, I would guess.
better than Opus? not even close. after struggling thru server overload for the past couple hours i finally put 5.1 thru the paces and it's....okay. failed some simple stuff that Sonnet/Opus/Gemini didn't. failed it badly and repeatedly actually. this was in typescript, btw. not sure if i'll keep the subscription or not
I appreciate that it's not working for your use case but it's unfortunate that you dismiss the experience of others. And i am not chinese, I am European. Thanks for your feedback anyway.
I tried Gemini 3.1 pro once to implement a previously designed 7-phase plan.
it only implemented a quarter of the plan before stopping, the code didnt even compile because half of the scaffolding was missing. it then confidently said everything was done.
Codex and GLM didnt have any issue following the exact same plan and getting a working app. So I would argue Gemini is the failure here.
I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition.
My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a flamewar or something.
It’s not a flame war, and you’re not just sharing your experience and encouraging others to try it out.
You’re making a claim, and I’m pointing out that it’s unsubstantiated and not consistent with any other source of data, including that internal to the company that makes the model.
I hope you can see that that’s different than saying it’s worked well for me
Sometimes we STEM folks are way too rigid, I obviously meant "IN MY OPINION, GLM models are at this point superior to...".
I do not think that anyone who read my comment understood it differently. But I grant you this point, this is just my opinion based on my personal experience not the result of a scientific study.
Once this is said, i wasn't submitting a scientific paper for preprint, just posting my opinion on an internet forum.
Not sure why you are making such a big deal out of it, especially for something for which people can decide within minutes if it works for them or not. And I haven't seen you nitpick on other people saying that all Chinese models are garbage incapable of doing even the most basic task, without quoting any study. This kind of scrutiny tends to be one-sided.
Edit: and regarding what the z.ai team is saying about their models, just check their Discord and the articles they link there. They themselves say that their latest models have leading performance on a number of aspects. It is misleading to suggest that the authors of the model are not proudly saying that their models have best in class performance.
Except it does not work that way in the US, you can freely incorporate in any state without worrying about this kind of tax drama. The EU really needs to improve the integration of their single market, as this is precisely the kind of barrier preventing people from exploring what other EU states have to offer.
reply