INsVI makes its research agenda visible while the work is underway. These inquiries guide experiments, dataset design, critical writing, and interdisciplinary collaboration.

Captioned datasets preserve what can be named, but often discard composition, atmosphere, material process, visual tension, ambiguity, and culturally specific ways of seeing.
What representations can retain relationships that ordinary labels flatten?
How can visual understanding be evaluated without turning every answer back into language?
What provenance and cultural information must remain attached to an image?
We distinguish generating a plausible image from understanding spatial structure, physical relations, compositional decisions, and changes across time.
Reasoning about depth, occlusion, scale, viewpoint, orientation, and navigation.
Recognizing hierarchy, rhythm, balance, contrast, and deliberate visual organization.
Understanding persistence, transformation, sequence, and causality across visual observations.
These are active research questions—not announcements of completed studies. INsVI will distinguish clearly among hypotheses, experiments, working papers, datasets, and peer-reviewed outcomes.