Approaches to streamlining and automating grounding?

I’ve been messing with media generation models a lot lately and noticed that MinimaxH3 seems to give surprisingly good quality ref2va video outputs even at low resolutions, when doing a character swap (cohesion, lack of significant artifacts, etc). That is, when i give a reference video and a reference image, and ask to swap the character. To me it seems like the takeaway is that it seems like the more rich and high quality the context provided is, it provides more “grounding” and “structure” for the model to provide a good output. In contrast to having to create a completely original scene just based on references, where it seems to get “lost in the woods” of the possibilities in latent space, and then the output has artifacts, distortions, inconsistencies. I would assume that this is a general principle across all models, even for text, audio, image, and video models. It would be like asking an LLM “make me rich” resulting in super generic ineffective advice, versus providing a rich, detailed, informative prompt regarding a current business plan and then asking a specific concrete question that maybe would help “narrow the latent space” and lead to a higher quality more useful output. Is there a name for this phenomenon? And what strategies do people use to help combat this issue, in ways that are more automated and don’t require painstaking effort, time, and labor crafting perfect prompts & context? An approach that is generalizable across tasks and requests, for example whether its a throwaway prompt like asking about travel itineraries, or a random prompt about a meme video generation, or even more detailed things about how to extract concrete actionable insights from a pool of business data. Hope this makes sense, still trying to wrap my mind around how to characterize this and articulate it better

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论