It seems likely to me this was driven by the `ultra` mode in 5.6, which fans subagents to do work. This mode was previously only available in the web UI (what was previously known as pro?)
It seems possible they trained this by doing full RL rollouts of agents interacting with each other. They likely view these prompts somewhat the same as raw reasoning traces, they don't want people to train directly on them.
I am unsure if this has been confirmed, but there are some signs that the opaque "compaction blob" they return from their dedicated compaction endpoint might not be text at all, rather a latent space representation of the conversation. The fact that OpenAIs compaction seems to be much higher fidelity than a lot of other providers makes me inclined to believe this.
If this is true, it doesn't seem far fetched to infer that they might be applying similar techniques to prompting subagents.
I would be curious to see if this way of spawning subagents (encrypted blob) is used when subagents of a different model type is spawned.
"Latent space representation" I have been waiting for this moment in the evolution of AI. Well, waiting with some trepidation. It seems inevitable that frontier AI's will, at some point, leave behind human-comprehensible representations of language. Purely for functional reasons, it's going to start making sense for AI agents to communicate amongst themselves in much more efficient ways than borrowing the languages of flesh-bag humans as an interface medium.
I Imagine next that programming languages, interfaces and API design starts going this direction next. Being written, expressed and optimized as blobs of high dimensional vector space. As humans we might still be able to understand some abstractions of what our AI's are talking about to each other, but maybe not more so then we understand how different regions of our own brain communicate with each other.
I strongly believe that the future is the other way. New programming languages and environments designed for strong auditability and preventing bugs will dominate. Only bad actors will use latent space representation, and it might even be outlawed. But the bad actors will proliferate underground…
Even for like token efficiency it could make sense - like imagine if the representation were more compact
Agents acn already translate languages quite well. It doesn't seem crazy that they could work and think in a model specific language, and then translate back to English or something for the user
It seems like the most efficient method would be for LLMs to communicated by exchanged latent space representations directly. Serial language is a incredibly inefficient way to encode these, a lot like flattening a complex graph into text.
I think you hit the nail on the head here. Having subagent dispatch in the loop for RLVR is something we've already seen in open models, like Kimi K2.5 and later, so it's no great stretch to assume OpenAI are doing it too.
If you keep RL'ing the dispatch then the prompts are likely to keep diverging from the type of prompt a person would write (like CoT becoming increasingly incomprehensible), and that divergence is part of their competitive advantage.
> rather a latent space representation of the conversation
That’s actually what got me to switch and use Codex sometime beginning this year, the compaction via these encrypted blobs was just waaaay better than Claude. I had short convos with Claude where it would forget something very obvious and important few million tokens into a task, whereas I reached ~1B tokens in some local codex sessions and it was recalling and paying attention to things I mentioned way back at the beginning of the session (and not persisted anywhere else in the repo/md files etc)
> It seems possible they trained this by doing full RL rollouts of agents interacting with each other. They likely view these prompts somewhat the same as raw reasoning traces, they don't want people to train directly on them.
this tracks. anthropic protects these as well iirc.
> I am unsure if this has been confirmed, but there are some signs that the opaque "compaction blob" they return from their dedicated compaction endpoint might not be text at all, rather a latent space representation of the conversation.
probably not a latent (to my knowledge latents aren't really part of the outer loop in ar-transformer inference processes), but maybe non-human-readable reasoning traces as occurs in fable.
They are not really token-in token-out per se, they are embedding-in embedding-out.
When operating on text, you embed each token into the LLMs embedding space. You go from a discrete token to a point in embedding space.
Likewise, when processing images, you have a image embedding model which produces a set of embedding vectors representing the contents of the image in the LLMs embedding (latent) space.
This same concept can be extended to compaction. Instead if limiting yourself to discrete tokens, you could generate a set of embedding vectors which represent the contents of the compacted conversation in latent space.
These have the possibility of containing a lot more semantic information per vector, which is why this can be appealing.
A big downside is decreased interpretability. AI safety people are generally fairly opposed to latent space reasoning for example, it can be harder to tell what the model is actually doing and if it is trying to deceive you.
If there is no visible prompt at all, then that is very understandable. The PR issue exposes a real gap though: subagent spawns need a human-readable audit trial, of its goals/intent, its boundaries and scope and limitations, etc; for basic responsible agentic harness functionality.
Add? Just make the sub-agents input prompt not encrypted, change "encrypted: true" to "encrypted: false" everywhere and everything continues to work as it used to (simplified, but you get the idea).
They need to fix the regression, not add something new here.
You can trivially enforce that at the AI provider level, which covers 99% of the problem the law is designed to address.
Of course it doesn't cover the issue of foreign state psyop operations but the fact that enforcing laws against organized crime and adversary state actors is hard isn't specific to AI.
Are you not aware of open-weights models and local generation? I think the vast majority of deepfake content is being genned in basements on RTX cards, not on public providers. People already have all this content, and have archives of it, and can run it airgapped. Cat is out of bag.
I would be very surprised if that would be the case. Maybe you mean deepfake content generated by organized crime or state actors, but that surely is a tiny fraction of what's being generated on Grok or other platforms.
I am well aware of them, and I'm well aware that they are very niche as I'm the only one of my surrounding to use one of those. And those very models are being developed by tech giants and VC backed companies, on which regulation have leverage.
The fact that a small black market exists doesn't mean regulating the mainstream market doesn't matters.
Also, most people like you fail to realizes that the EU only has mandate from the member states to regulate the economy. The EU has no business dealing with people using SDXL finetunes on RTX cards in their garage.
No. Again, this regulation is about regulating businesses because that's what the EU is about.
The general use or creation of deepfake for porn, harassment, or election manipulation, is outside of what the EU can regulate as an institution, it is the responsibility of member states. (The same way the EU can impose rules on platform with respect to copyright violations, but cannot enact rules against piracy in general, these are always made by member states).
You don't have to prove anything? You just have to mark the outputs of your slop generator appropriately. "Proving" one way or another is their problem when it comes to enforcement.
What makes it so great for me is the effortlessness.
I often use Python for quick one off scripts. With UV I can just do `uv init`, `uv add` to add dependencies, and `uv run` whatever script I am working on. I am up and running in under a minute. I also feel confident that the setup isn't going to randomly break in a few weeks.
With most other solutions I have tried in the Python ecosystem, it always seemed significantly more brittle. It felt more like a collection of hacks than anything else.
Yeah, the endorsement was mentioned - while the assassination attempt was not.
Here is a quote from that study:
> In Phase Two, comparing Republican-leaning and Democrat-leaning accounts, we again observed an engagement shift around the same date, affecting all metrics
It's so obvious that an assassination attempt on a republican candidate will boost republican engagement MORE than democrats. It's just as obvious that a corona virus outbreak in close proximity to a laboratory which works on corona viruses, escaped very likely from that lab. Common sense. Also for weeks after the event, of course republicans continued to be more riled up about it.
This is not science. If the claim is that Musk tuned the algorithms at that special date, you have to prove it by other means, other than engagement being boosted at that day and weeks to follow.
It seems to me that when taken together, these two make a fair case that there was a change in algorithm on the given date?
Besides, the study you are claiming "is not science" did not even make strong claims as to there being an algorithm change. The study made an analysis, and concluded that they might point towards there being an algorithmic change.
Additionally, as you say "If the claim is that Musk tuned the algorithms at that special date, you have to prove it by other means". What other means exactly? We have no data to go on except observations of what happens on the platform. From the information we have it seems like everything points towards this being the case.
The case for manipulation happening furthered even more by Elon Musk repeatedly showing he is full of shit, and has no qualms lying and totally making stuff up. (see his repeated lies about Autopilot, his lie about being world class at video games, his lies about DOGE cuts, etc).
Before dismissing these findings, ask yourself honestly: would you apply the same rigorous standards of proof if the algorithmic changes benefited different political figures or viewpoints?
> would you apply the same rigorous standards of proof if the algorithmic changes benefited different political figures or viewpoints?
That's the whole point I'm basically making: I think this is an obviously bogus study, but one side will happely take the conclusions as granted because it falls in line with an agenda or narrative - furthering the divisions in society. Which by the way doesn't mean that the other side doesn't spew garbage as well.
But in these specific case, where it's so obvious... We can go on with corona virus lab like I mentioned. In these cases, where the common sense is "turned off" in some people, I really wonder what's going on and speak up. It's absurd.
> The study made an analysis, and concluded that they might point towards there being an algorithmic change.
Yes but you can't do that for the reasons I mentioned. But still doing it and then another study referencing that crap, shows to me that science is not at play here.
> From the information we have it seems like everything points towards this being the case.
Well if Greta Thunberg at the height of her popularity fell from a wind mill, someone could've made the claim Twitter is suddenly boosting a certain group of accounts. The CEO of Twitter even wrote condolences, something is up here!! Sorry, but this garbage.
> What other means exactly? We have no data to go on except observations of what happens on the platform.
Yes that's an issue, but that's not my problem?! If the circumstances do not allow for a proper study, you don't make it. I would look whether the boost in republican engagement came down again. If it stayed on the same level (maybe til today?!), I would agree, something is very fishy. Maybe there is a study that did that already? I don't know.
> The case for manipulation happening furthered even more by Elon Musk repeatedly showing he is full of shit, and has no qualms lying and totally making stuff up.
But this is not a science approach you can enrich a crappy study with. You can surely have that opinion - I have no problem with that. That's an opinion without proof, which I have too on certain topics. But then producing bogus studies to try to turn this opinion into some sort of fact where people point to as "proof" is what makes me upset. It doesn't help the cause, only causes division, because the people pointing think they have science on their side, instead of just having a opinion.
Your dismissal of this research shows a fundamental misunderstanding of how scientific analysis works.
The study analyzed two separate phenomena - changes in engagement for Musk's posts AND changes in engagement for Republican content - occurring simultaneously. When taken together, these patterns strongly suggest algorithmic changes, not just organic user behavior from the assassination attempt.
Saying "you can't do that for the reasons I mentioned" is just wrong. Studies absolutely can point toward likely explanations without 100% certainty - that's literally how science progresses. The researchers used appropriate cautious language because they understand scientific rigor, not because their analysis is "crap."
Your argument that "if circumstances don't allow for a proper study, you don't make it" would eliminate most scientific advancement. Should we have abandoned Alzheimer's research because perfect data wasn't available? Obviously not.
What's truly absurd here is your selective skepticism. You demand impossibly high standards of proof for findings you dislike while accepting "common sense" explanations that align with your preconceptions.
Before dismissing research as "garbage" that "causes division," maybe consider whether your reaction is based on methodological concerns or simply that the evidence contradicts your preferred narrative about Musk. Your eagerness to defend him while offering nothing but personal opinion suggests it's the latter.
> The study analyzed two separate phenomena - changes in engagement for Musk's posts AND changes in engagement for Republican content - occurring simultaneously
And in which political bubble is Musk popular? It's not seperate phenomena. What do you expect? That let's say republican engagement organically increases (as I claim), but Musks engagement stays the same, even though he made many statements regarding the assassination attempt and got people riled up because of his endorsement? This is ridiculous, sorry.
> Should we have abandoned Alzheimer's research because perfect data wasn't available? Obviously not.
The study authors should've been well aware of the above. They didn't care and did it anyway, because they knew fully well that people will not care, because it fits a certain narrative.
> Your eagerness to defend him
Okay I think we can stop it here then. You think I write in defense of Musk. I think you are completely oblivious to common sense. The truth is probably not that black and white though.
We were talking about his son, he let the process play out in full which is beyond respectable.
When the incoming administration has shown extraordinary will to persecute political opponents, I think it would be unethical not to preemptively pardon these people.
Yeah, i feel like currently they are at about the price of camera traps 10 years ago. There is very little mass-manufacturability to them right now (it's all open source and made from off-the-shelf parts) but later if we can find more funding, we are going to make a design more for manufacturing which should hopefully drive the costs down even more! :)
> There is very little mass-manufacturability to them right now (it's all open source and made from off-the-shelf parts)
This is the obstruction to using them in an educational setting. If they were available for $600+ each but already completely built (minimal DIY), they would be more likely to get into (some) schools.
We have a group of kids in Rhode Island building some with the library there! Part of a "Wildlives" program where the kids also learn to put camera traps around the local nature!
totally! Right now we are just trying to get them out and tested on science projects around the world, but hopefully we can find funding to make more designs that could be manufactured in bulk (like the audiomoth and groupgets) and have even more of these things out and about!
It is now. There were a few years where it had basically disappeared (2015-2018). When Apple eventually put it back in the open-source world, it was done with little fanfare so it could be easy to miss.
Helpful thread, thanks: Google support team churn after the distribution transition to Asus IoT, Frigate devs were preparing to fork Google repos, then new Google devs appeared.
> Google is getting back on top of things aka coral support which is nice.. it seems that the original devs weren't on the project and new devs needed to be given notice. Hopefully this continues and things are kept up to date.. updated libcoral and pycoral libraries are coming as well.
It's good that Frigate brought attention to languishing Linux maintenance for Coral. Rockchip 3588 and other Arm SoCs have NPUs, which will likely be supported in time, but each SoC will require validation. Coral Edge TPUs were a convenient single target that worked with any x86 and Arm board, via USB or M.2 slot.
It seems possible they trained this by doing full RL rollouts of agents interacting with each other. They likely view these prompts somewhat the same as raw reasoning traces, they don't want people to train directly on them.
I am unsure if this has been confirmed, but there are some signs that the opaque "compaction blob" they return from their dedicated compaction endpoint might not be text at all, rather a latent space representation of the conversation. The fact that OpenAIs compaction seems to be much higher fidelity than a lot of other providers makes me inclined to believe this.
If this is true, it doesn't seem far fetched to infer that they might be applying similar techniques to prompting subagents.
I would be curious to see if this way of spawning subagents (encrypted blob) is used when subagents of a different model type is spawned.