Hacker Newsnew | past | comments | ask | show | jobs | submit | bluejay2387's commentslogin

I think a lot of were already using Mermaid and/or Python/Matplotlib(etc) for this. What would be the advantages to Flint?

Mostly for expressiveness and reliability & cost trade-off. Flint has advantage of being an intermediate language that allow agents to generate good-looking stuff without additional refinement loops, since the compiler derives lower-level geometric constraints from semantic types.

Would especially be handy when we are building some agents that produce charts that serve end users (they want faster and more reliable experiences!).


"End times are nothing new, it's the historic default mode." -- might be the smartest thing I have read in many weeks.


The entire domain of NP-Complete problems would beg to differ with you.


I'm curious how good AI is at these.

We have plenty of them that humans even find to be pleasant to solve, such as Minesweeper :)


About 90% of my coding is on Qwen 3.6 27b and Open Code with some custom skills and Semble. It is NOT as smart as CC or Codex but its enough to get most of my work done. I didn't set out to replace CC and Codex (I have an RTX 6000 so the TPS is faster than I care about, but the RTX 6000 was originally for other work). I only tried this just to see how close you could get to a frontier model for coding as an experiment, but it was good enough that I stuck with it. I still fall back to Codex for really complicated stuff and to polish UI's as that seems to be the weakest element to working in Qwen.This isn't a recommendation because I don't think most people have an RTX 6000 laying around and the cost would be many years of MAX CC or Codex subscriptions, but at least this seems possible. Maybe in a few more years it will even be practical.

Other Notes: I have had to set the compact target to 75% on a 256k context window as once the conversation length goes about 100k I start seeing a drop in the quality and speed. This becomes very problematic after about 150k. I tried Qwen 3.5 122b too but it actually seems much worse at coding than 3.6 27b even though its much larger. Maybe because I am using a 4bit quant or maybe I just don't have it configured correctly? I know 3.6 is newer but I didn't expect it to out perform a model that is much larger from the prior generation. Gemma 4 31b is a good model for other tasks but at least my personal experience is that Qwen outperforms in coding. Nemotron Super 120b is great at a lot of stuff but it also seems to be not as good at coding as Qwen. This was very surprising to me.


Same here, I use Qwen 3.6 27b (Q6 quant) with llama.cpp on an RTX 5090 using the pi agent exclusively now. The fact that it's local means that I never have to think about token pricing, quotas, time of day, or data sensitivity. I have limited the GPU from 600W to 450W which means the system stays whisper quiet during inference.

I have become so "lazy" (in a good way), so far that I've started using the model for lots of daily mundane things on top of just coding:

  * "commit this on a branch, push, create a PR and assign $nickname for review"
  * "Use the Stripe CLI to download all open and overdue invoices and reconcile them with this CSV export from our bank account."
  * "Use these Elasticsearch credentials to summarise what kind of operations are causing load at the moment."
  * "Tell me if our codebase already supports X and where it's  implemented."


What context length and kv cache quant (if any) are you using? And MTP?


No KV cache quant, context length 50% of original, MTP absolutely. These are the relevant cmdline attributes. Getting around 100t/s with this setup, even when watt-limited to 450W.

  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --metrics --jinja --chat-template-file chat_template.jinja --chat-template-kwargs '{"preserve_thinking": true}' --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75 -ngl 99 -c 131072 -fa on -np 1 -hf unsloth/Qwen3.6-27B-MTP-GGUF:Q6_K


Not the person you asked, but I have a 9700 which has the same VRAM, and running Q6 on it with unquantized kv gives me 50k context. Putting -ctv q8_0 ups that to 70k. I normally run Q4 with unquantized kv @ 130k at 50 t/s (mtp 3), with the disclaimer that I'm running PCIe gen4x8, so I'm slightly slowed. I've found that quantizing k leads to broken json on tool calls, which is fairly unrecoverable, but YMMV.


Qwen3.5-122B is actually Qwen3.5-122B-A10B. The A10B means that this is a "mixture of experts" model where only 10B parameters are activated at a given time. Whereas Qwen3.6-27B is a "dense" model where all 27B parameters are activated all the time. So for many tasks, you'd expect the 27B dense model to be better than the 122B-A10B model.


I am forced to use Qwen 3.6 27b at work and found it next to useless. I might as well do all the work manually rather than having it implement another mess or get the debugging entirely wrong.

It feels like anything less than Sonnet is just a waste of time, apart from use as a smarter search function.

It also strikes me as strange that you would mention Codex for UI polish, as it's notoriously bad at UI, and far behind Claude Opus. Altman specifically posted that they are working to improve this for the next model release.


It might be good at analysis & review, writing documentation, git commits, etc--even if it's not good at coding.

All the drudgery.


Bad AI written documentation and commits are not great, particularly when you work in a team.

I almost find it offensive when colleagues open a MR with an obvious slop description that's frequently inaccurate.

That said, I find AI useful for a lot of drudgery like resolving merge conflicts or splitting changes out into separate MRs.

Particularly with the latter I had issues with small models, they butchered the changes I wanted moved. Not even on the second attempt did GPT 5.4 mini manage to move 10-20 lines to another file without modifying them in the process.


why 27b vs 35b? Is MoE that much worse for coding?


Can take the geometric mean of total and active parameters of MoE to get approximate equivalent quality to dense model params. So sqrt(35*10)≈18.7.

The trade-off of MoE is that it is worse but faster for the same total size.


Yeah MoE is a little worse for the same size, but you can often run bigger MoEs at respectable speeds even on cpu ram offload. The dense models really need to be 100% vram


I have a very similar setup: OpenCode, Qwen3.6-27B (llama-server with an RTX 5090). It works well for my purposes. As a semi coding luddite, I use it for mundane tasks.


In the US -- once our nation finishes attacking our own education system -- this is definitely something a group of academic institutions could get together and accomplish. I assume the same is true in other countries. Companies like Nvidia and AMD might even support that effort, as they make money on the hardware and would probably be more than happy for there to be more reasons to use it. There may have not been a compelling enough motivation to achieve this before, but "models" didn't have this level of strategic relevance until relatively recently. Nvidia has been fairly good about releasing open weight models in the last few months.


Wait, which side is blocking kids fork taking algebra or forcing universities to admit people that can't do math or read, or abandoning phonetics for unproven methods that don't work?



Both sides, since they are bought and paid for by the finance industrial complex.


It's the US, both "sides" of that coin are bad with examples pro and con all over the shop.

Still, to specifically give a partial answer to your poor faith rhetorical just askin' musing: Florida Conservatives

(specifically turfing nerds from New College of Florida and bringing an excess number of baseball sports bro's to a place that likes math and has no baseball field)


From what I can tell the majority of developers here have moved into the "Anger" stage.


I had a locally hosted model write its own semantic search system that indexed 250,000 documentation and code files and then write a fully functioning mod for one of the games I play based on that documentation that I couldn't get to work after 2 weeks of my own effort, all in under 4 hours (and that included a 25 minute long indexing process). This freaked me out enough that I then had it write a CLI based activity and TODO tracker and then integrate that tool into its coding process to track all of its activities in about another 2 hours. I am still emotionally recovering from this day. I have since replaced the semantic search system with an open source option (though I used it for a few months) but I still use the activity tracker for both coding projects and myself.


What mod did you build?


A mod that fixed a bug that prevented certain buffs from working when mounted for the Magus class / Arcane Rider archetype in Pathfinder Wrath of the Righteous. It also managed to fix the problem with Shelters not providing protection from corruption when resting in outposts in that same mod. I've used other models to expand the mod to an entire mini-expansion with new Archetypes and abilities since then.


I am a 'fan' of Open Web UI, but the document editing mode is a compelling feature that Open Web UI does not have. I'll probably wait a while before trying Odysseus... let the inevitable security problems work themselves out.


In a related story... I got led on by Eliza. I tried to have a productive conversation and she just kept asking me redundant questions. It's obvious that she was trying to extend the conversation for nefarious reasons that I can only guess at. It's true I approached her and started the conversation, but I hardly think that makes me blamable for what happened here.


I’m sorry you feel that way — can you tell me more about what made you feel led on?


Yes. Yes it does. Eliza is a known AI. You choose to expose yourself to its output. You are 100% culpable for your actions that sprang from your interactions.


Did you forget the /s ?


I have exposure to AI initiatives at several companies including a few F500's. I have seen teams dump huge logs into frontier models that took hours to get so-so results that we were able to replace with a few lines of python code at 1000 times the speed and 100% accuracy. When asked why they were doing this they literally said "because we don't understand the subject matter so we were depending on the AI". I saw one team file a complaint with a vendor about a frontier backed coding harness and it's inability to consistently format headers because they were using it as a reporting engine. When I recommended they just use the coding tool to write code to generate reports you would have thought I had just cured cancer from their response. I frequently see people complain about the fact that AI is going to take their jobs and then see them gripe about the fact that AI is 'worthless' because it can't do more of their job than it already does. It's easy to see the difference between the people seeing 10x productivity gains from leveraging AI and those who aren't and it's not the AI.


i have trouble understanding these situations, e.g. the AI itself would presumably make the suggestion to write a python script for such a task. It seems to me that there two huge problems right now * understanding which category of problems an LLM is an appropriate solution for (rather than throwing LLMs at any and all problems) * matching model capability (and therefore cost) to the problem at hand. You can easily overspend massively by using a model that's too powerful


Someone asked me if I was using models for fantasy sports, and if it was smart enough to help make decisions about drafting.

My answer: no, but it was able to help me find the website and social handles for every beat writer for every team, and generate a simple website where I can do a daily skim of teams/players and draw my own conclusions.

LLMs are a tool, not a panacea.


I've heard this framed as "AI raises the floor by 2x or less but raises the ceiling by 10x or more"


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: