Qwen3.8-Omni-Flash: One AI for text, images, audio, and video
Qwen3.8-Omni-Flash(Qwen three point eight Omni Flash)
An AI model that accepts text, images, audio, and video.
multimodal(mul-tee-MOH-dul)
Able to handle several kinds of information together.
context window(CON-text WIN-doh)
The amount of information an AI can consider in one request.
What happened
The Qwen team at Alibaba released Qwen3.8-Omni-Flash. It accepts text, images, audio, and video. Qwen's official documentation lists audio and video analysis, meeting summaries, and subtitle generation. It also lists a 1M-token context window. The model is presented as turning those inputs into text. Qwen's official documentation
The story also appeared on Hacker News. The post received 327 points and 126 comments. Those numbers show community attention. They do not prove the announcement or the model's quality. Hacker News post
The background
A video carries several kinds of information at once. Pictures show actions. Audio carries speech and other sounds. Text may appear on the screen. Time connects these pieces. A multimodal model tries to handle them together.
That matters because a long recording can lose meaning when cut into small pieces. A large context window can let a model consider more of the recording at once. It is not permanent memory. It is the amount of information available during one request.
Why it matters
The main promise is not only that the model can see and hear. It can connect mixed inputs to practical tasks. A meeting assistant could use a recording to create notes. A subtitle workflow could use speech and video context together. These are possible workflows, not guarantees.
The official model materials also list tool calling and web search. Tool calling lets a model ask another program to perform a task. Web search can provide outside information. Together, they point toward workflows where an AI analyzes media, looks something up, and returns a text result. Accuracy still depends on the model and the connected tools.
What is confirmed
Qwen's documents list Chat Completions and Responses as access routes. They also list support for tool calling and web search. Qwen's launch materials report a more than 26% average gain across 30 evaluations, compared with Qwen3.5-Omni-Plus. That is a claim from the provider's evaluation. It is not the same as an independent replication. Qwen launch post
On Hacker News, the item received 327 points and 126 comments. This is useful evidence of interest, not evidence that the model works as claimed.
What remains unknown
We still need independent tests on long videos, noisy audio, accents, and mixed languages. We also need real measurements for price, speed, and error rates. A 1M-token limit does not tell us how much usable information the model can accurately follow.
What to watch next
Watch for third-party comparisons, real API costs, and examples with clear error checks. The careful conclusion is modest: Qwen3.8-Omni-Flash adds a broad input option for media-heavy work. It does not yet prove that one model understands every long video perfectly.
Qwen3.8-Omni-Flash: An AI for words, pictures, sound, and video
📰 Full story: Qwen3.8-Omni-Flash: One AI for text, images, audio, and video
A new AI model can work with text, images, audio, and video.
context window(CON-text WIN-doh)
The information an AI can consider at one time.
tool calling(tool CALL-ing)
A way for an AI to ask another program for help.
Hacker News(HACK-er News)
A technology news discussion site.
💡 The gist
- It can take text, images, audio, and video.
- It can help summarize meetings and make subtitles.
- Hacker News attention shows interest, not proof.
What is new?
Qwen3.8-Omni-Flash is a new model from Alibaba's Qwen team. It can take several kinds of information in one request. The official documents list audio and video analysis. They also mention meeting summaries and subtitle generation. Qwen's official documentation
The model has a 1M-token context window. A context window is the information an AI can consider. A bigger window can help with long recordings. It may help the model connect a decision with details mentioned later. It is not the same as permanent memory.
The model also supports tool calling and web search. Tool calling lets an AI ask another program for help. Web search lets it find information online. These tools can build larger workflows. However, tools do not guarantee correct answers.
How should we read the numbers?
Qwen reports a gain of more than 26% across 30 evaluations. The comparison used the earlier Qwen3.5-Omni-Plus model. That result comes from Qwen's own report. Independent tests could produce different results. Qwen launch post
The Hacker News post had 327 points and 126 comments. Those numbers measure attention from that community. They do not test the model themselves. They also do not prove the model is correct.
What comes next?
We still need tests with long videos and noisy audio. We need real prices, speeds, and error rates. We also need to see how well the model handles different languages. Next, watch independent comparisons and careful examples. Qwen3.8-Omni-Flash looks useful for media-heavy work. It is not proof of perfect understanding.
💬 Qwen 3.8 Omni Flash: strong promise, but hard to run locally
Commenters liked the reported ability and low price. However, most real-world speed and bug reports are about related Flash Next/Max models, not necessarily Omni Flash itself.
- A commenter said the article reports Omni Flash as close to Gemini 3.8 Flash for audio and video, and possibly better at audio. This has not been independently checked in the thread.
- The commenter-supplied price comparison is $0.15 input/$0.47 output for Qwen and $1.50/$9.00 for Gemini per million tokens. A cheaper token does not always mean a cheaper completed task.
- Users discussing the related 125B Flash Next model report that it needs about 128 GB of RAM. Estimates put a 40+ tok/s setup near $3,000 and a 4090 setup near 30 tok/s; the 27B model is easier to run.
- Some users saw strange reasoning, made-up context, or very long reasoning. Serving and quantization may be involved; one user reported improvement with NVFP4, but the issue is not confirmed as universal.
- People disagree about whether open-weight releases are slowing down: Qwen 3.8 Max and Flash Next are cited as counterexamples.
mature digest at 137 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.
Qwen3.8-Omni-Flash can look at many kinds of things
📰 Full story: Qwen3.8-Omni-Flash: One AI for text, images, audio, and video
This AI can take words, pictures, sounds, and videos.
Qwen3.8-Omni-Flash(Qwen three point eight Omni Flash)
An AI made by Alibaba's Qwen team.
Hacker News(HACK-er News)
A website where people discuss technology news.
What is it?
Qwen3.8-Omni-Flash is an AI from Alibaba's Qwen team. It can take words, pictures, sounds, and videos. Then it gives an answer with words.
What can it do?
It can help make short notes from a meeting. It can help make subtitles for a video. It can look at a lot of information together. It can also ask computer tools for help.
Why are people talking?
People talked about this news on Hacker News. The post had 327 points and 126 comments. Those numbers show attention from readers. They do not prove the AI is correct.
What should we remember?
AI can still make mistakes. People should check important answers. We still need more tests with long videos.
💬 Qwen 3.8 Omni Flash: clever, but big and sometimes quirky
The comments suggest real promise, but the reports were not all made under the same conditions.
- A commenter said the article makes Omni Flash sound close to Gemini 3.8 Flash for sound and pictures, and possibly stronger for sound. That has not been independently confirmed.
- User-provided prices are $0.15/$0.47 per million input/output tokens for Qwen and $1.50/$9.00 for Gemini. But a task can use many tokens, so the cheaper number may not mean a cheaper job.
- The related big 125B model may need about 128 GB of computer memory. Users estimate about $3,000 for more than 40 tok/s and about 30 tok/s on a 4090; the smaller 27B model is easier.
- Some users saw strange thoughts, made-up context, or long delays, possibly because of how the model was run or compressed. Others still found it useful. People also disagree about whether open-weight releases are slowing, with Max and Flash Next cited as released.
mature digest at 137 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.
💬 Qwen 3.8 Omni Flash: promising capability, costly local execution
HN commenters were excited by a reported audio/video comparison and low prices, but most hands-on speed and reliability reports concern related Qwen 3.8 Flash Next/Max models, not necessarily Omni Flash.
mature digest at 137 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.