AI agents are becoming increasingly capable of reasoning, planning, and interacting with external tools, but moving from a successful prototype to a reliable production system remains a significant engineering challenge. This session explores practical techniques for engineering and evaluating AI agents that perform consistently in real-world environments.
Discussion includes how to design agentic workflows, define meaningful evaluation metrics, and measure performance beyond simple accuracy. Topics include task completion, reasoning quality, tool selection, LLM-as-a-Judge, automated evaluation pipelines, human evaluation, regression testing, and production monitoring.