Pada si Iroyin
ỌjaAI Understanding finifini

MarkTechPost ṣe ijabọ Google imudojuiwọn Gemini Omni 1.1 Filaṣi pẹlu awọn amugbooro fidio to gun ati awọn iṣakoso fireemu

MarkTechPost ṣe ijabọ pe Google's Gemini Omni 1.1 Filaṣi ṣe afikun to awọn aaya 10 ti ipo ipo, akopọ 40-keji, iṣakoso akọkọ-ati-kẹhin-fireemu, din owo 360p iyaworan, ati 4K upscaling. Ijabọ naa sọ pe awoṣe wa nipasẹ Google's Gemini API, AI Studio, ati Platform Aṣoju Idawọlẹ Gemini.

7 min readRead the linked source
Source-provided image accompanying MarkTechPost reports Google updated Gemini Omni 1.1 Flash with longer video extensions and frame controls
itọkasi orisunOrisun ti o gbasilẹ
Olutẹwe
marktechpost.com
Orisun ọna asopọ
marktechpost.comhttps://www.marktechpost.com/2026/08/29/google-ai-releases-gemini-omni-1-1-flash-40-second-scene-extension-first-last-frame-control-and-4k-upscaling/
Orisun iru
Orisun ti o sopọ mọ - ipo orisun akọkọ ko ti fi idi mulẹ.
Tun toka si

Ìtàn gbẹyìn tunwo

AtokọLoye eyi ni iṣẹju 60

Bẹrẹ nibi

Awọn ofin bọtini

API (Àwòrán Ètò Ìlò)
Ọna ti a ṣeto fun eto sọfitiwia kan lati firanṣẹ awọn ibeere si ati gba awọn idahun lati eto miiran.
Iwọn otutu
Eto iṣapẹẹrẹ ti n ṣakoso aileto ni awọn abajade ti ipilẹṣẹ.
Agbara
Agbara awoṣe lati ṣetọju iṣẹ ṣiṣe labẹ ariwo, awọn iyipada, tabi awọn igbewọle ọta.
Ṣe idanwo fun ara rẹKini AI? Idanwo

Ohun ti yi pada niwon atejade

  1. Ni akọkọ ti a tẹjade
  2. This documentation is a material technical update to the same Gemini Omni 1.1 Flash release covered by the canonical entry. It adds the model ID, preview status, release date, input and output modalities, video-generation and editing functions, sound-generation support, token and file limits, supported formats, regional information and consumption restrictions.
  3. This Google DeepMind blog materially advances the continuing Gemini Omni 1.1 Flash release by detailing scene extensions that use up to 10 seconds of prior context and reach 40 seconds cumulatively, first-and-last-frame interpolation, 360p previews that Google claims are up to 60% faster and one-third the cost of 720p, 1080p and 4K upscaling, video-reference inputs of up to three seconds, and availability through Google AI Studio, the Agent Platform API, Google Flow and the Gemini app.
  4. MarkTechPost materially advances the continuing Gemini Omni 1.1 Flash release story by reporting implementation, availability, pricing, resolution, provenance, and editing-limit details, including 10 seconds of scene context, 40-second cumulative extensions, first-and-last-frame control, and 360p draft rendering. These details are attributed to MarkTechPost and are not independently confirmed in the supplied source.
  5. This candidate materially advances the same Gemini Omni 1.1 Flash release already represented by the eligible canonical update. MarkTechPost reports additional concrete details about the update’s 10-second scene context, cumulative 40-second extension ceiling, first-and-last-frame interpolation, video-reference limits, 360p draft economics, SynthID watermarking, regional restrictions, unsupported controls, pricing, and named production users. Those details are attributed to MarkTechPost and have not been independently confirmed here.

Kini o ṣẹlẹ

MarkTechPost reports that Google released Gemini Omni 1.1 Flash, an update to its multimodal video-generation and editing model. The reported changes focus on directability and iterative editing: the model can analyze up to 10 seconds of prior video when extending a scene, accept pinned first and last frames, use short video references for character or object consistency, and preserve editing state across conversational turns through the Interactions API. MarkTechPost says the model is available through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow. The outlet says Adobe, Figma Weave, GMI Cloud, and Runway are already production users, but those customer-use claims have not been independently confirmed here.

MarkTechPost reports that Google released Gemini Omni 1.1 Flash as a production update to its native multimodal video-generation and editing model. According to the outlet, it accepts text, images, audio, and video in one workflow and produces video with audio. Stateful conversational editing runs through Google’s Interactions API: passing a previous_interaction_id lets the model retain earlier video state and apply a named change without uploading the prior video again. The source illustrates this by making a generated violin invisible while preserving other unspecified elements, but does not independently establish how consistently that preservation works.

Scene extension reportedly uses up to 10 seconds of prior context, compared with only the final second in earlier versions. The model generates a three-to-10-second continuation, with extensions made in 10-second increments up to a cumulative 40-second ceiling. MarkTechPost says the added context is intended to preserve motion, characters, and audio, and that some final frames may be edited to make the seam continuous. Extensions append only to the end, not the beginning or middle. Uploaded inputs must be no longer than 10 seconds unless a model-generated video is extended through a multi-turn interaction. New dialogue cannot be added when extending an uploaded clip in which someone is already speaking, while spoken dialogue can be supported in multi-turn extension.

The update also reportedly provides first-and-last-frame interpolation: developers supply a starting frame and ending frame, and the model generates motion between them. MarkTechPost identifies camera orbits, dolly-zooms, and seamless loops as intended uses, without independent demonstrations or measurements. Prompt tags can assign media roles, including a still image as a style or subject reference and a video as a character or object reference. Video references are limited to three clips of up to three seconds each; audio inside a video reference is ignored, and reasoning across multiple videos is unsupported and may reduce output quality. The model supports 360p, 720p, 1080p, and 4K, with 720p default and the two higher resolutions produced through upscaling. Google reportedly says 360p previews generate up to 60% faster and cost one-third as much as 720p. Reported pricing is $1.50 per 1 million input tokens, $9 per 1 million output text tokens, and $17.50 per 1 million output video tokens, with an effective cost of about $0.10 per second for 720p under standard pricing. Availability through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow has not been independently verified.

Awọn alaye orisun: marktechpost.com ↗

Kini idi ti o ṣe pataki

The reported update addresses practical bottlenecks in AI video production: maintaining continuity, controlling transitions, and iterating without regenerating an entire clip. A draft-at-360p workflow could reduce the cost and time of experimentation, while 1080p and 4K outputs are available as upscaled resolutions. These capabilities may matter to developers and creative teams building video tools, although the source does not provide independent testing of output quality, reliability, latency, or customer results.

If the reported controls work reliably, Gemini Omni 1.1 Flash could move some AI-video workflows from repeated full regeneration toward targeted editing. Preserving unmentioned elements across conversational turns could reduce prompt restatement, clip reconstruction, and manual combination of generations, while longer scene context could reduce discontinuities. First-and-last-frame control adds explicit boundary conditions: a starting frame establishes where a clip begins and an ending frame constrains where it finishes, potentially helping transitions, loops, and controlled camera movement. However, the report does not establish physical plausibility, subject consistency, edit accuracy, temporal consistency, identity preservation, or audio continuity through neutral testing.

The reported draft-resolution option may make experimentation more accessible to teams paying per generated output. MarkTechPost says 360p previews are faster and substantially cheaper than 720p, while 1080p and 4K are upscaled outputs. A draft-then-upscale workflow could lower prompt-iteration costs and help teams reject weak concepts before a final render. The source does not compare fine detail, text rendering, faces, motion artifacts, or audio quality between draft and final outputs. The model is also paid-only and has no provisioned-throughput option, which may limit predictable large-scale use. These capabilities may matter to developers and creative teams building video tools, but the source provides no independent testing of output quality, reliability, latency, or customer results.

The reported provenance feature may matter to publishers and platforms handling synthetic media. MarkTechPost says every generated video carries an invisible SynthID watermark that viewers cannot see but software can detect. That could support downstream identification, but the source does not explain detection coverage, after editing or compression, access to detection tools, or whether the watermark identifies the particular model or generation; it should not be treated as a complete authenticity or attribution system. The named production customers—Adobe, Figma Weave, GMI Cloud, and Runway—could signal early adoption, but the report does not specify deployment scope, features, volumes, performance results, or independent confirmation. This review treats those statements as claims attributed to MarkTechPost rather than established facts.

Interactive Mechanism

Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ

Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Ibanisọrọ Erongba Ṣayẹwo+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

Kini lati wo tókàn

The main questions are whether the reported controls produce dependable continuity outside carefully prepared examples and how the model performs on dialogue, multiple subjects, complex motion, and longer workflows. The source identifies significant limits: extensions append only to the end, uploaded clips must generally be 10 seconds or shorter, voice editing and audio references are unsupported, and uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. MarkTechPost says its claims were checked against Google documentation and pricing, but this review has not independently confirmed those documents, prices, availability, or customer deployments.

The first issue is real-world continuity over repeated extensions. The source says the model can use up to 10 seconds of context and extend a clip to 40 seconds, but reports no success rates or independent tests. Future evaluations should examine characters, objects, camera motion, lighting, spatial relationships, and audio across successive extensions, distinguishing model-generated clips from uploaded clips because their reported rules differ. Independent testing should also examine difficult transitions in which frames differ in pose, viewpoint, lighting, object arrangement, or scene geometry. Although the source names orbits, dolly-zooms, and seamless loops, it does not establish performance, failed generations, unwanted subject changes, temporal artifacts, or required prompt iteration.

Availability and regional restrictions will affect use. MarkTechPost reports availability through several Google services, but says uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. It also says there is no free tier and no provisioned throughput. Developers will need to confirm current access, regional eligibility, quotas, pricing, and delivery requirements for files above 4MB. The source says outputs above that size require URI delivery and polling through the Files API until the file becomes active. The unsupported controls are also consequential: the model reportedly lacks system instructions, , top_p, stop sequences, negative-prompt controls, voice editing, audio references, and YouTube URLs as source inputs. Negatives must be placed in ordinary prompt text.

Language coverage and provenance require follow-up. MarkTechPost says English is fully supported while other languages are unevaluated. The most useful evidence would include reproducible tests, current API documentation, clear regional terms, and independent users describing actual workflows. MarkTechPost says its claims were checked against Gemini API documentation and pricing, but this review has not independently confirmed Google’s documentation, the stated prices, listed availability, SynthID behavior, or named production deployments. Customer statements should describe actual use rather than simply naming companies as users, and testing should cover dialogue, multiple subjects, complex motion, and longer workflows.

Awọn itọsọna ti o jọmọ & awọn ibeere

Kini AI?Awọn awoṣe AI ti ṣalayeÌlànà Ìwà AIỌjọ́ Iwájú AIṢe idanwo ohun ti o mọ — gbiyanju idanwo AI ọfẹ kanWa ọrọ AI kan ninu iwe-itumọ waTẹle olutọpa idasilẹ awoṣe AI

Awọn imudojuiwọn ati awọn atunṣe

Itan alamọdaju yii ti ni imudojuiwọn ni aye nigbati iṣẹlẹ to sese ndagbasoke nipa ti ara. URL rẹ ati ọjọ ikede atilẹba ko yipada.

  • This candidate materially advances the same Gemini Omni 1.1 Flash release already represented by the eligible canonical update. MarkTechPost reports additional concrete details about the update’s 10-second scene context, cumulative 40-second extension ceiling, first-and-last-frame interpolation, video-reference limits, 360p draft economics, SynthID watermarking, regional restrictions, unsupported controls, pricing, and named production users. Those details are attributed to MarkTechPost and have not been independently confirmed here.
  • MarkTechPost materially advances the continuing Gemini Omni 1.1 Flash release story by reporting implementation, availability, pricing, resolution, provenance, and editing-limit details, including 10 seconds of scene context, 40-second cumulative extensions, first-and-last-frame control, and 360p draft rendering. These details are attributed to MarkTechPost and are not independently confirmed in the supplied source.
  • This Google DeepMind blog materially advances the continuing Gemini Omni 1.1 Flash release by detailing scene extensions that use up to 10 seconds of prior context and reach 40 seconds cumulatively, first-and-last-frame interpolation, 360p previews that Google claims are up to 60% faster and one-third the cost of 720p, 1080p and 4K upscaling, video-reference inputs of up to three seconds, and availability through Google AI Studio, the Agent Platform API, Google Flow and the Gemini app.
  • This documentation is a material technical update to the same Gemini Omni 1.1 Flash release covered by the canonical entry. It adds the model ID, preview status, release date, input and output modalities, video-generation and editing functions, sound-generation support, token and file limits, supported formats, regional information and consumption restrictions.
Wo awọn àkọsílẹ awọn atunṣe log
Ṣe eyi wulo?