Google's Gemini Enterprise Agent Platform has introduced implicit and explicit context caching for its Gemini models. This feature is designed to reduce costs and latency for AI requests containing repeated content, offering up to a 90% discount on cached tokens for certain models. This is particularly useful for scenarios such as chatbots and repetitive analysis of large video or document files.