If you've played around with 'raw' LLM interactions you've already seen this: feeding a prompt into an LLM which ends with a 'start of user' prompt will produce a plausible query into the agent. Which makes perfect sense because the LLMs are already trained on many examples of this and the 'predict the next token' loss function does not particularly distinguish between the sides of the conversation. I highly doubt they need this feature to get better training data, more likely they got this feature for free from the way that the training works and only recently decided to actually expose it to the user.
I work at a company that has an enterprise contact that should mean I'm not part of the training data. But as a manual mode power user I think about this.
I'd love the open source community to come up with a way to harvest model usage by experienced software engineers before we forget our crafts. I'm not against auto mode, but it's not something a couple private companies should have monopolies on.
I've always hated interfaces that try to complete my sentences for me. It started with suggested replies in email and IM apps. I sure noticed it when they started showing up in the llm chat interfaces and it really bugs me.
This is the one case I don't mind it. Suggested responses feel like they cheapen human interaction, but here I'm talking to a robot that really does tend to know what I want next, and also is highly unlikely to be offended by a less-than-heartfelt response.
The suggestions in the box where I put my text push me out of flow. Bad enough it says “want to do X?”, but when it’s in my own text box it screws me up. It’s not more efficient. It actively makes things worse for me.
The worst is Gmail's recent feature to suggest an entire goddamn email that includes cheery little details, doing its best to mimic human pleasantries and small talk. Rather than simply offering "yep/nope" type short replies like it once did, it'll now auto-compose and suggest a multi-paragraph email responding to questions like "How's the family doing?" or "Is your older cat tolerating the new kitten yet?" with completely fabricated saccharine slop.
It's like Clippy pops up and goes "It looks like you're trying to maintain a shred of human connection in an online interaction. Would you like a smiling skinwalker to do that for you instead?"
This is hilarious as someone who hasn't used the gmail composition interface in years. I left it to use my own domain, but it seems like I get to enjoy this mess with popcorn too instead of tears.
How does showing the suggested answers to the user make the conversation better for model training?
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
The idea is that a thread can have many reasonable follow-ups that the user would've accepted, so it is wrong to punish the model for predicting a follow up that is different from the user message, as that prediction could've been accepted by the user if it was given.
RL training, the second phase of LLM training, is based on "I did X, was that good/bad?" and that 1 bit of information is the training data.
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.
I started getting prompts about "how is claude doing?" as a separate thing in Claude Code, that I noticed yesterday. So they're (also?) soliciting direct feedback about satisfaction with the session.
And if you do provide feedback, they also collect the session. So it's a way for them to collect prompts, answers and overall grade for how good the answers are.
This isn't really convincing, since you can do this even without showing the prediction at all. Simply ask the model to predict what the user will send, then show the actual next prompt, and done. The only reason to show this would be to influence the user's next prompt, which the article doesn't touch on.
This has been around for so long that I was using it for a while and then stopped (a while back) when I moved to using just the mobile app (and even it shows suggestions, but I have no idea how to invoke) for everything.
One annoyance I have is the suggested prompt is not a bad idea, but not what I want to do next. But it interrupts me and sometimes I go with it. So I don't think it's a accurate prediction, more like a self-fulfilling prophecy.
I'd love the open source community to come up with a way to harvest model usage by experienced software engineers before we forget our crafts. I'm not against auto mode, but it's not something a couple private companies should have monopolies on.
It's like Clippy pops up and goes "It looks like you're trying to maintain a shred of human connection in an online interaction. Would you like a smiling skinwalker to do that for you instead?"
The iOS keyboard and gmail app are the worst offenders here.
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.
Interesting thought at least.