This week, Microsoft released two new in-house AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, into public preview. These models are designed to enhance Microsoft's AI capabilities across their product suite, reducing reliance on third-party models like those from OpenAI.
MAI-Image-2.5-Pro is Microsoft's most advanced image generation model to date. It is targeted at high-fidelity applications such as hero imagery and detailed editing, with particular improvements in precise in-image text rendering. The model is priced at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. It has already gained recognition in the creative industry, with advertising executives noting its potential for generative media tools.
On the other hand, MAI-Voice-2-Flash is optimized for speed and cost-efficiency, running twice as fast as its predecessor, MAI-Voice-2, and costing 32% less. Priced at $15 per million characters, it is designed for high-volume voice applications such as call centers and real-time speech operations, where cost and latency are critical factors.
These models showcase Microsoft's strategy of developing a diverse family of AI models rather than a single flagship, catering to different user needs. MAI-Image-2.5-Pro is already integrated into Bing, PowerPoint, and OneDrive, where it has demonstrated significant cost and efficiency improvements. MAI-Voice-2-Flash is utilized in Dynamics 365 Contact Center and Azure Voice Live, with notable GPU cost reductions reported.
The models are particularly aimed at enterprises and developers. MAI-Image-2.5-Pro suits creative industries needing high-quality image generation, while MAI-Voice-2-Flash addresses the needs of sectors requiring efficient voice processing, such as customer service operations.
Work implications: These models could enhance productivity in creative and customer service roles by providing more efficient tools for image and voice processing.
Originally reported by VentureBeat