DeepSeek recently released a new experimental multimodal model, claiming it will bring performance close to Anthropic’s Opus 4.8 model on matters of agentic capabilities.

    DeepSeek-V4-Flash-Vision-Exp is available through the Chinese AI developer’s API, adding image interpretation abilities to its V4-Flash model. The system can process images alongside text, allowing it to interpret screenshots, charts and other visual information.

    DeepSeek said the model “matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge” and makes “a major leap” over V4-Flash on multimodal agent benchmarks. The company added its model’s multimodal agent performance is “close to Opus-4.8”.

    The new model supports JPEG, PNG, GIF and WebP images and can process up to 600 images in a single request, according to DeepSeek. The company also released an update to its DeepSeek Harness tool with support for the new model.

    DeepSeek has previously developed multimodal models, including its DeepSeek-VL family.

    Source: Mobile World Live

    Image Credit: DeepSeek




    Source: Tahawul Tech