ok sorry - can you explain more why it'd be a game changer to stream the tool call. Whats an example tool call that this would help with. Most of the time the tools I want called are at the end of the conversation (i.e. "summarize what we talked about and email it to me"). Or if its in the middle of the conversation, we can just carry on the conversation while the tool call is happenning in parallel..
"It'd be a game-changer to be able to have the model start replying with partial information streamed from the tool call, then seamlessly continue with additional information."
You can't continue a conversation if it hinges on information that takes inherently takes 30+ seconds to generate. Previously even models that "handle it" can only play for time.
You can try and simulate this by messing with model context mid-conversation but that breaks down reliability massively as the model loses track of what it's talking about.