Hi J,
Thanks for the reply. A little more info:
In my specific case, I am indeed using guided generation already (which helped a lot in seed 1 and 2).
I thought the eval session was great. I've been working with LLMs for a bit now and it's awesome to have a whole WWDC session to talk about evals and expose them to devs who may not have seen them before.
This is coming up for me now because something changed in seed 3 where my feature went from a ~95% success rate to a 0% success rate, all failing with guardrails errors that did not trigger in the first two seeds. There are a bunch of other threads here on the topic and I've filed several feedbacks on the specifics already.
Maybe that's a bug/unexpected outcome and we'll see a future seed restore the behavior. I hope so, I'd like to ship this feature. But if not, at least I won't have sent it to customers.
My real concern is that in a future point release where the amount of feedback time is compressed and IMO it's very hard to get specific issues in front of engineers via the Feedback system with enough time to get a fix done and tested, this will pop up again. Since the models are non-deterministic, unless you're running my evals, there's a good chance the team may not even know about what's for me a serious regression until it ships.
I really want to use and love the framework because it's so promising but I'm a little freaked out now given what's happening in the current seeds. That's the background for my concern and I don't think there's any way to mitigate that at the moment (right?)