Topic: flash
18 stories found
Yesterday
Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88%
Google introduced Agentic Video Understanding in Gemini flash models, enabling the platform to navigate videos efficiently rather than processing them at 1 FPS, thereby reducing video tokens by up to 88% and improving performance. This update is significant as it enhances Gemini's efficiency and responsiveness when handling video content.
Friday, September 4, 2026
ggml/llama.cpp releases: v0.4.0
Version 0.4.0 of llama.cpp was released, adding support for Qwen3.8-Flash-Next and Nemotron-3-Puzzle models, along with several new features like on-demand tensor reading and video input options, making it more versatile for AI language tasks. This update is significant as it enhances the model's capabilities and flexibility, catering to a broader range of applications in natural language processing.
Wednesday, September 2, 2026
Monday, August 31, 2026
Saturday, August 29, 2026
Friday, August 28, 2026
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture
ggml/llama.cpp releases: b10665
The ggml/llama.cpp project added DSpark support for Nemotron3.5 in version b10665, enhancing model compatibility and functionality. This update is significant as it broadens the software's applicability for users with specific hardware configurations.
Thursday, August 27, 2026
Gemini Omni 1.1 Flash lets you build with more control
Gemini Omni 1.1 Flash software update enhances building controls, allowing for greater customization and functionality. This upgrade is significant as it empowers architects and engineers to design more sophisticated structures with improved precision and flexibility.
Wednesday, August 26, 2026

Qwen3.8-Flash-Next
Qwen3.8, an updated version of the Qwen AI model, was released with enhanced features to improve natural language processing capabilities. This update is significant as it aims to provide more accurate and contextually relevant responses, potentially advancing the state of AI in text generation and understanding.
Ollama releases: v0.33.1
Ollama released version 0.33.1, which includes updates to Qwen3.8 Flash Next support, cmake patches, and mlxrunner structured output for improved model loading times, highlighting ongoing development and community contributions.
18 of 18 items shown. Sources: 107 days indexed.