Visual bridge for text-only models and visual enhancement toolkit for multimodal AI agents — OCR, zoom, PDF/Office rendering, web vision and interaction.