Question 274
A developer builds an MCP tool that generates images using Stable Diffusion. The tool takes a prompt and returns a base64-encoded image. Claude then describes the image to the user. However, Claude's descriptions don't match the generated images because Claude can't actually see the tool output as an image. How should this be architected?