Kini o ṣẹlẹ
MarkTechPost reports that Google released Gemini Omni 1.1 Flash, an update to its multimodal video-generation and editing model. The reported changes focus on directability and iterative editing: the model can analyze up to 10 seconds of prior video when extending a scene, accept pinned first and last frames, use short video references for character or object consistency, and preserve editing state across conversational turns through the Interactions API. MarkTechPost says the model is available through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow. The outlet says Adobe, Figma Weave, GMI Cloud, and Runway are already production users, but those customer-use claims have not been independently confirmed here.
MarkTechPost reports that Google released Gemini Omni 1.1 Flash as a production update to its native multimodal video-generation and editing model. According to the outlet, it accepts text, images, audio, and video in one workflow and produces video with audio. Stateful conversational editing runs through Google’s Interactions API: passing a previous_interaction_id lets the model retain earlier video state and apply a named change without uploading the prior video again. The source illustrates this by making a generated violin invisible while preserving other unspecified elements, but does not independently establish how consistently that preservation works.
Scene extension reportedly uses up to 10 seconds of prior context, compared with only the final second in earlier versions. The model generates a three-to-10-second continuation, with extensions made in 10-second increments up to a cumulative 40-second ceiling. MarkTechPost says the added context is intended to preserve motion, characters, and audio, and that some final frames may be edited to make the seam continuous. Extensions append only to the end, not the beginning or middle. Uploaded inputs must be no longer than 10 seconds unless a model-generated video is extended through a multi-turn interaction. New dialogue cannot be added when extending an uploaded clip in which someone is already speaking, while spoken dialogue can be supported in multi-turn extension.
The update also reportedly provides first-and-last-frame interpolation: developers supply a starting frame and ending frame, and the model generates motion between them. MarkTechPost identifies camera orbits, dolly-zooms, and seamless loops as intended uses, without independent demonstrations or measurements. Prompt tags can assign media roles, including a still image as a style or subject reference and a video as a character or object reference. Video references are limited to three clips of up to three seconds each; audio inside a video reference is ignored, and reasoning across multiple videos is unsupported and may reduce output quality. The model supports 360p, 720p, 1080p, and 4K, with 720p default and the two higher resolutions produced through upscaling. Google reportedly says 360p previews generate up to 60% faster and cost one-third as much as 720p. Reported pricing is $1.50 per 1 million input tokens, $9 per 1 million output text tokens, and $17.50 per 1 million output video tokens, with an effective cost of about $0.10 per second for 720p under standard pricing. Availability through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow has not been independently verified.
Awọn alaye orisun: marktechpost.com ↗
Kini idi ti o ṣe pataki
The reported update addresses practical bottlenecks in AI video production: maintaining continuity, controlling transitions, and iterating without regenerating an entire clip. A draft-at-360p workflow could reduce the cost and time of experimentation, while 1080p and 4K outputs are available as upscaled resolutions. These capabilities may matter to developers and creative teams building video tools, although the source does not provide independent testing of output quality, reliability, latency, or customer results.
If the reported controls work reliably, Gemini Omni 1.1 Flash could move some AI-video workflows from repeated full regeneration toward targeted editing. Preserving unmentioned elements across conversational turns could reduce prompt restatement, clip reconstruction, and manual combination of generations, while longer scene context could reduce discontinuities. First-and-last-frame control adds explicit boundary conditions: a starting frame establishes where a clip begins and an ending frame constrains where it finishes, potentially helping transitions, loops, and controlled camera movement. However, the report does not establish physical plausibility, subject consistency, edit accuracy, temporal consistency, identity preservation, or audio continuity through neutral testing.
The reported draft-resolution option may make experimentation more accessible to teams paying per generated output. MarkTechPost says 360p previews are faster and substantially cheaper than 720p, while 1080p and 4K are upscaled outputs. A draft-then-upscale workflow could lower prompt-iteration costs and help teams reject weak concepts before a final render. The source does not compare fine detail, text rendering, faces, motion artifacts, or audio quality between draft and final outputs. The model is also paid-only and has no provisioned-throughput option, which may limit predictable large-scale use. These capabilities may matter to developers and creative teams building video tools, but the source provides no independent testing of output quality, reliability, latency, or customer results.
The reported provenance feature may matter to publishers and platforms handling synthetic media. MarkTechPost says every generated video carries an invisible SynthID watermark that viewers cannot see but software can detect. That could support downstream identification, but the source does not explain detection coverage, after editing or compression, access to detection tools, or whether the watermark identifies the particular model or generation; it should not be treated as a complete authenticity or attribution system. The named production customers—Adobe, Figma Weave, GMI Cloud, and Runway—could signal early adoption, but the report does not specify deployment scope, features, volumes, performance results, or independent confirmation. This review treats those statements as claims attributed to MarkTechPost rather than established facts.
Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ
Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.
A route planner searches possible journeys using explicit rules. What does this illustrate about AI?
Kini lati wo tókàn
The main questions are whether the reported controls produce dependable continuity outside carefully prepared examples and how the model performs on dialogue, multiple subjects, complex motion, and longer workflows. The source identifies significant limits: extensions append only to the end, uploaded clips must generally be 10 seconds or shorter, voice editing and audio references are unsupported, and uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. MarkTechPost says its claims were checked against Google documentation and pricing, but this review has not independently confirmed those documents, prices, availability, or customer deployments.
The first issue is real-world continuity over repeated extensions. The source says the model can use up to 10 seconds of context and extend a clip to 40 seconds, but reports no success rates or independent tests. Future evaluations should examine characters, objects, camera motion, lighting, spatial relationships, and audio across successive extensions, distinguishing model-generated clips from uploaded clips because their reported rules differ. Independent testing should also examine difficult transitions in which frames differ in pose, viewpoint, lighting, object arrangement, or scene geometry. Although the source names orbits, dolly-zooms, and seamless loops, it does not establish performance, failed generations, unwanted subject changes, temporal artifacts, or required prompt iteration.
Availability and regional restrictions will affect use. MarkTechPost reports availability through several Google services, but says uploaded-video editing or extension is unavailable in the EEA, Switzerland, and the United Kingdom. It also says there is no free tier and no provisioned throughput. Developers will need to confirm current access, regional eligibility, quotas, pricing, and delivery requirements for files above 4MB. The source says outputs above that size require URI delivery and polling through the Files API until the file becomes active. The unsupported controls are also consequential: the model reportedly lacks system instructions, , top_p, stop sequences, negative-prompt controls, voice editing, audio references, and YouTube URLs as source inputs. Negatives must be placed in ordinary prompt text.
Language coverage and provenance require follow-up. MarkTechPost says English is fully supported while other languages are unevaluated. The most useful evidence would include reproducible tests, current API documentation, clear regional terms, and independent users describing actual workflows. MarkTechPost says its claims were checked against Gemini API documentation and pricing, but this review has not independently confirmed Google’s documentation, the stated prices, listed availability, SynthID behavior, or named production deployments. Customer statements should describe actual use rather than simply naming companies as users, and testing should cover dialogue, multiple subjects, complex motion, and longer workflows.