Institute for Visual Intelligence · New York

Intelligence
begins with
seeing.

A research institute investigating visual intelligence beyond the limits of language.

INsVI brings fine arts, machine perception, spatial reasoning, and data science into one field of inquiry.

Abstract visual study for Research at INsVI

What we investigate

01 / PERCEPTION

Visual
intelligence

How machines form spatial, compositional, and perceptual understanding—not merely captions about images.

Abstract visual study for Current Inquiries at INsVI
02 / MATERIAL

Visual data
as culture

How datasets encode artistic practice, cultural judgment, absence, bias, and ways of seeing.

Abstract visual study for Visual Datasets at INsVI
03 / PRACTICE

Art after
generative AI

How artists can test, critique, and reshape intelligent systems through visual experimentation.

Abstract visual study for Art and Machine Perception at INsVI
Our premise

Images are not
sentences in disguise.

The institute

Language models describe the visible world. Visual intelligence must learn its structure, relations, ambiguities, and force.

Current inquiry / 2026

Visual Intelligence Beyond the LLM

Large language models can name, caption, and discuss images, but that is not the same as seeing as an artist or designer sees. INsVI studies visual intelligence as a distinct mode of thought exercised through making, selecting, comparing, composing, revising, and judging visual form.

This is not simply spatial intelligence. Spatial intelligence concerns location, orientation, geometry, navigation, and objects in space. Artistic and design intelligence also works through proportion, rhythm, contrast, hierarchy, color interaction, material behavior, ambiguity, metaphor, cultural context, affect, and aesthetic judgment.

Read the current inquiry

Beyond language-mediated vision

Multimodal systems often privilege what can be verbalized. They may recognize what is depicted while failing to understand why a composition works, which subtle alteration weakens it, or how an artist discovers meaning by looking and making before that thought can be stated in words.

LLMs and visual evidence

Fluent language can conceal weak visual understanding. Captions, labels, embeddings, and preference scores can flatten texture, scale, sequence, context, and relations among parts. For AGI, the question is not merely whether a system can talk about the visible world, but whether it can form, test, and revise visual judgments without treating language as the sole measure of understanding.

The dataset is part of the intelligence

Web-scale datasets contain duplication, poor reproductions, missing provenance, inaccurate captions, cultural imbalance, uncertain rights, and images detached from their material setting. A serious art-and-design dataset should preserve authorship, provenance, medium, dimensions, viewing conditions, process, revision, relationships among works, interpretation, and disagreement. Dataset construction is a curatorial and epistemological practice—not merely an engineering task.

What INsVI is seeking

We seek to articulate the visual intelligence used by artists and designers: the capacity to discover meaning through form, organize attention, perceive relationships before they become propositions, and make judgments whose evidence remains substantially visual.

Read the full chronology and inquiry →

Abstract visual study for About INsVI at INsVI