Model recommender, Prompt Optimizer, and Prompt benchmark for System Prompts
I’ve tried the Agent Architect but didn’t have much luck with getting a good answer. My question related to what reasoning level and model I should use to process my payload to avoid deadline errors and it recommended GPT4o which had a harder time processing the information I was sending.
On the prompt optimizer I was thinking of something similar to OpenAI’s prompt optimizer but trained for compatibility with your Agent Studio environment and documentation for System Prompts.
It would also be useful to have a way to benchmark a prompt to quickly see the output without having to re-trigger another webhook payload. e.g. An option to replay a recent Webhook payload from recent logs.