Skip to content

I have a question about maskgit cross-attention with text tokens #8

Description

@9B8DY6

In cvivit, you would train video clip with the fixed number of frames.
When training maskgit to do cross attention with text tokens, how did you cut(?) corresponding text tokens for the given frames?
Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions