Retour aux Actualités
ProduitBriefing AI Understanding

MarkTechPost rapporte que Google a mis à jour Gemini Omni 1.1 Flash avec des extensions vidéo plus longues et des contrôles d'image

MarkTechPost rapporte que le flash Gemini Omni 1.1 de Google ajoute jusqu'à 10 secondes de contexte de scène, des extensions cumulées de 40 secondes, un contrôle de la première et de la dernière image, des brouillons 360p moins chers et une mise à l'échelle 4K. Le rapport indique que le modèle est disponible via l'API Gemini de Google, AI Studio et la plate-forme d'agent d'entreprise Gemini.

7 min readRead the linked source
Source-provided image accompanying MarkTechPost reports Google updated Gemini Omni 1.1 Flash with longer video extensions and frame controls
Référence sourceSource enregistrée
Éditeur
marktechpost.com
Lien source
marktechpost.comhttps://www.marktechpost.com/2026/08/29/google-ai-releases-gemini-omni-1-1-flash-40-second-scene-extension-first-last-frame-control-and-4k-upscaling/
Type de source
Source liée : le statut de source principale n'a pas été établi.
Également cité

Histoire révisée pour la dernière fois

ContexteComprenez cela en 60 secondes

Commencez ici

Termes clés

API (interface de programmation d'applications)
Une manière structurée permettant à un système logiciel d'envoyer des requêtes et de recevoir des réponses d'un autre système.
Température
Un paramètre d'échantillonnage contrôlant le caractère aléatoire des sorties générées.
Robustesse
Capacité d'un modèle à maintenir ses performances malgré le bruit, les changements ou les entrées contradictoires.
Testez-vousQu’est-ce que l’IA ? Quiz

Ce qui a changé depuis la publication

  1. Première publication
  2. This documentation is a material technical update to the same Gemini Omni 1.1 Flash release covered by the canonical entry. It adds the model ID, preview status, release date, input and output modalities, video-generation and editing functions, sound-generation support, token and file limits, supported formats, regional information and consumption restrictions.
  3. This Google DeepMind blog materially advances the continuing Gemini Omni 1.1 Flash release by detailing scene extensions that use up to 10 seconds of prior context and reach 40 seconds cumulatively, first-and-last-frame interpolation, 360p previews that Google claims are up to 60% faster and one-third the cost of 720p, 1080p and 4K upscaling, video-reference inputs of up to three seconds, and availability through Google AI Studio, the Agent Platform API, Google Flow and the Gemini app.
  4. MarkTechPost materially advances the continuing Gemini Omni 1.1 Flash release story by reporting implementation, availability, pricing, resolution, provenance, and editing-limit details, including 10 seconds of scene context, 40-second cumulative extensions, first-and-last-frame control, and 360p draft rendering. These details are attributed to MarkTechPost and are not independently confirmed in the supplied source.
  5. This candidate materially advances the same Gemini Omni 1.1 Flash release already represented by the eligible canonical update. MarkTechPost reports additional concrete details about the update’s 10-second scene context, cumulative 40-second extension ceiling, first-and-last-frame interpolation, video-reference limits, 360p draft economics, SynthID watermarking, regional restrictions, unsupported controls, pricing, and named production users. Those details are attributed to MarkTechPost and have not been independently confirmed here.

Que s'est-il passé

MarkTechPost reports that Google released Gemini Omni 1.1 Flash, an update to its multimodal video-generation and editing model. The reported changes focus on directability and iterative editing: the model can analyze up to 10 seconds of prior video when extending a scene, accept pinned first and last frames, use short video references for character or object consistency, and preserve editing state across conversational turns through the Interactions API. MarkTechPost says the model is available through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow. The outlet says Adobe, Figma Weave, GMI Cloud, and Runway are already production users, but those customer-use claims have not been independently confirmed here.

MarkTechPost reports that Google released Gemini Omni 1.1 Flash as a production update to its native multimodal video-generation and editing model. According to the outlet, it accepts text, images, audio, and video in one workflow and produces video with audio. Stateful conversational editing runs through Google’s Interactions API: passing a previous_interaction_id lets the model retain earlier video state and apply a named change without uploading the prior video again. The source illustrates this by making a generated violin invisible while preserving other unspecified elements, but does not independently establish how consistently that preservation works.

Scene extension reportedly uses up to 10 seconds of prior context, compared with only the final second in earlier versions. The model generates a three-to-10-second continuation, with extensions made in 10-second increments up to a cumulative 40-second ceiling. MarkTechPost says the added context is intended to preserve motion, characters, and audio, and that some final frames may be edited to make the seam continuous. Extensions append only to the end, not the beginning or middle. Uploaded inputs must be no longer than 10 seconds unless a model-generated video is extended through a multi-turn interaction. New dialogue cannot be added when extending an uploaded clip in which someone is already speaking, while spoken dialogue can be supported in multi-turn extension.

The update also reportedly provides first-and-last-frame interpolation: developers supply a starting frame and ending frame, and the model generates motion between them. MarkTechPost identifies camera orbits, dolly-zooms, and seamless loops as intended uses, without independent demonstrations or measurements. Prompt tags can assign media roles, including a still image as a style or subject reference and a video as a character or object reference. Video references are limited to three clips of up to three seconds each; audio inside a video reference is ignored, and reasoning across multiple videos is unsupported and may reduce output quality. The model supports 360p, 720p, 1080p, and 4K, with 720p default and the two higher resolutions produced through upscaling. Google reportedly says 360p previews generate up to 60% faster and cost one-third as much as 720p. Reported pricing is $1.50 per 1 million input tokens, $9 per 1 million output text tokens, and $17.50 per 1 million output video tokens, with an effective cost of about $0.10 per second for 720p under standard pricing. Availability through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow has not been independently verified.

Détails de la source: marktechpost.com ↗

Pourquoi c'est important

The reported update addresses practical bottlenecks in AI video production: maintaining continuity, controlling transitions, and iterating without regenerating an entire clip. A draft-at-360p workflow could reduce the cost and time of experimentation, while 1080p and 4K outputs are available as upscaled resolutions. These capabilities may matter to developers and creative teams building video tools, although the source does not provide independent testing of output quality, reliability, latency, or customer results.

If the reported controls work reliably, Gemini Omni 1.1 Flash could move some AI-video workflows from repeated full regeneration toward targeted editing. Preserving unmentioned elements across conversational turns could reduce prompt restatement, clip reconstruction, and manual combination of generations, while longer scene context could reduce discontinuities. First-and-last-frame control adds explicit boundary conditions: a starting frame establishes where a clip begins and an ending frame constrains where it finishes, potentially helping transitions, loops, and controlled camera movement. However, the report does not establish physical plausibility, subject consistency, edit accuracy, temporal consistency, identity preservation, or audio continuity through neutral testing.

The reported draft-resolution option may make experimentation more accessible to teams paying per generated output. MarkTechPost says 360p previews are faster and substantially cheaper than 720p, while 1080p and 4K are upscaled outputs. A draft-then-upscale workflow could lower prompt-iteration costs and help teams reject weak concepts before a final render. The source does not compare fine detail, text rendering, faces, motion artifacts, or audio quality between draft and final outputs. The model is also paid-only and has no provisioned-throughput option, which may limit predictable large-scale use. These capabilities may matter to developers and creative teams building video tools, but the source provides no independent testing of output quality, reliability, latency, or customer results.

The reported provenance feature may matter to publishers and platforms handling synthetic media. MarkTechPost says every generated video carries an invisible SynthID watermark that viewers cannot see but software can detect. That could support downstream identification, but the source does not explain detection coverage, after editing or compression, access to detection tools, or whether the watermark identifies the particular model or generation; it should not be treated as a complete authenticity or attribution system. The named production customers—Adobe, Figma Weave, GMI Cloud, and Runway—could signal early adoption, but the report does not specify deployment scope, features, volumes, performance results, or independent confirmation. This review treats those statements as claims attributed to MarkTechPost rather than established facts.

Interactive Mechanism

Mécanisme interactif : comment cela fonctionne réellement

Explorez de manière interactive la technologie sous-jacente à ce développement.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Vérification de concept interactive+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

Que regarder ensuite

The main questions are whether the reported controls produce dependable continuity outside carefully prepared examples and how the model performs on dialogue, multiple subjects, complex motion, and longer workflows. The source identifies significant limits: extensions append only to the end, uploaded clips must generally be 10 seconds or shorter, voice editing and audio references are unsupported, and uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. MarkTechPost says its claims were checked against Google documentation and pricing, but this review has not independently confirmed those documents, prices, availability, or customer deployments.

The first issue is real-world continuity over repeated extensions. The source says the model can use up to 10 seconds of context and extend a clip to 40 seconds, but reports no success rates or independent tests. Future evaluations should examine characters, objects, camera motion, lighting, spatial relationships, and audio across successive extensions, distinguishing model-generated clips from uploaded clips because their reported rules differ. Independent testing should also examine difficult transitions in which frames differ in pose, viewpoint, lighting, object arrangement, or scene geometry. Although the source names orbits, dolly-zooms, and seamless loops, it does not establish performance, failed generations, unwanted subject changes, temporal artifacts, or required prompt iteration.

Availability and regional restrictions will affect use. MarkTechPost reports availability through several Google services, but says uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. It also says there is no free tier and no provisioned throughput. Developers will need to confirm current access, regional eligibility, quotas, pricing, and delivery requirements for files above 4MB. The source says outputs above that size require URI delivery and polling through the Files API until the file becomes active. The unsupported controls are also consequential: the model reportedly lacks system instructions, , top_p, stop sequences, negative-prompt controls, voice editing, audio references, and YouTube URLs as source inputs. Negatives must be placed in ordinary prompt text.

Language coverage and provenance require follow-up. MarkTechPost says English is fully supported while other languages are unevaluated. The most useful evidence would include reproducible tests, current API documentation, clear regional terms, and independent users describing actual workflows. MarkTechPost says its claims were checked against Gemini API documentation and pricing, but this review has not independently confirmed Google’s documentation, the stated prices, listed availability, SynthID behavior, or named production deployments. Customer statements should describe actual use rather than simply naming companies as users, and testing should cover dialogue, multiple subjects, complex motion, and longer workflows.

Guides et quiz associés

Qu’est-ce que l’IA ?Modèles d'IA expliquésÉthique de l'IAAvenir de l'IATestez ce que vous savez : essayez un quiz gratuit sur l'IARecherchez un terme d'IA dans notre glossaireSuivez le suivi des versions du modèle AI

Mises à jour et corrections

Cette histoire canonique est mise à jour lorsque l’événement en développement change matériellement. Son URL et sa date de publication originale ne changent jamais.

  • This candidate materially advances the same Gemini Omni 1.1 Flash release already represented by the eligible canonical update. MarkTechPost reports additional concrete details about the update’s 10-second scene context, cumulative 40-second extension ceiling, first-and-last-frame interpolation, video-reference limits, 360p draft economics, SynthID watermarking, regional restrictions, unsupported controls, pricing, and named production users. Those details are attributed to MarkTechPost and have not been independently confirmed here.
  • MarkTechPost materially advances the continuing Gemini Omni 1.1 Flash release story by reporting implementation, availability, pricing, resolution, provenance, and editing-limit details, including 10 seconds of scene context, 40-second cumulative extensions, first-and-last-frame control, and 360p draft rendering. These details are attributed to MarkTechPost and are not independently confirmed in the supplied source.
  • This Google DeepMind blog materially advances the continuing Gemini Omni 1.1 Flash release by detailing scene extensions that use up to 10 seconds of prior context and reach 40 seconds cumulatively, first-and-last-frame interpolation, 360p previews that Google claims are up to 60% faster and one-third the cost of 720p, 1080p and 4K upscaling, video-reference inputs of up to three seconds, and availability through Google AI Studio, the Agent Platform API, Google Flow and the Gemini app.
  • This documentation is a material technical update to the same Gemini Omni 1.1 Flash release covered by the canonical entry. It adds the model ID, preview status, release date, input and output modalities, video-generation and editing functions, sound-generation support, token and file limits, supported formats, regional information and consumption restrictions.
Voir le journal des corrections publiques
Vous avez trouvé cela utile ?