Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Ling-3.0-flash-VL adds image and video understanding to Ling-3.0-flash, and the interesting part is that vision did not sit alongside the language model — it raised it. On the Artificial Analysis intelligence index the visual model scores 42 against 38 for the text model it was built from, on the same 124-billion-parameter foundation with only 5.5 billion active per inference step. A vision encoder feeds a two-layer projector that aligns visual features with text, while a space-and-time positional encoding handles video. What it was built to do is act on what it sees: read a design reference and write the interface, render the result and compare it against the original, identify elements on a screen and operate them.
