PROGRAMS & SELECTED ACTIVITIES

From the present to the beginning.

A descending chronology of inquiry, programs, field observation, and the development of visual intelligence as an institutional question.

Working chronology. Archival images are shown where verified. Additional annual fair documentation will be added as photographs and records are organized.

PRESENT

Visual Intelligence Beyond the LLM

2023–present · Current inquiry

Large language models now dominate artificial intelligence, but their fluency can obscure a basic limitation: being able to name, caption, or discuss an image is not the same as being able to see as an artist or designer sees. INsVI’s current inquiry examines visual intelligence as a distinct mode of thought—one exercised through making, selecting, comparing, composing, revising, and judging visual form.

This is not simply spatial intelligence. Spatial intelligence concerns location, orientation, geometry, navigation, depth, and the manipulation of objects in space. Artistic and design intelligence includes those capacities but extends to relationships that cannot be reduced to coordinates: proportion, rhythm, contrast, hierarchy, color interaction, material behavior, visual tension, ambiguity, metaphor, historical reference, cultural context, affect, and aesthetic judgment.

Read more about the current inquiry

Beyond language-mediated vision

Multimodal models commonly process images through representations optimized for prediction and alignment with language. They can recognize depicted entities and produce persuasive explanations, yet may still fail to understand why a composition works, which subtle alteration weakens it, how materials constrain a form, or why two visually similar works make different cultural and aesthetic claims. An artist or designer often reasons by looking and making before the thought can be fully stated in words.

INsVI therefore asks whether visual intelligence includes irreducibly nonverbal operations: sustained attention, perceptual comparison, sensitivity to negative space and edge relationships, recognition of visual consequence, tolerance of productive ambiguity, and iterative judgment through the emerging work itself. These capacities differ both from linguistic reasoning and from narrowly defined spatial problem-solving.

The limits of current LLM-centered systems

LLMs are exceptionally effective at synthesizing textual patterns. When attached to vision systems, however, they can privilege what is easy to verbalize. The conversion of an image into captions, labels, embeddings, or preference scores can flatten scale, texture, sequence, context, and relationships among parts. Confident language may conceal weak visual evidence, and consensus-driven prediction can suppress unfamiliar forms that do not fit established categories.

The question for AGI is therefore not whether a system can talk about the visible world, but whether it can form, test, and revise visual judgments without treating language as the sole measure of understanding.

The dataset is part of the intelligence

More images do not automatically produce better seeing. Web-scale datasets contain duplication, poor reproductions, missing provenance, inaccurate captions, cultural and geographic imbalance, uncertain rights, and images detached from their material scale or original setting. Preference datasets can reduce complex aesthetic responses to a binary ranking. As synthetic imagery re-enters training corpora, errors and stylistic averages may be recursively amplified.

For art and design, a serious dataset must preserve more than a final image. It should document authorship and provenance; date, place, medium, dimensions, and viewing conditions; process and revision; relationships among works; expert and community interpretations; and the uncertainty or disagreement surrounding aesthetic judgment. Dataset construction is therefore a curatorial, ethical, and epistemological practice—not merely an engineering task.

What INsVI is seeking

The aim is not to replace spatial intelligence or linguistic intelligence, but to articulate the visual intelligence used by artists and designers: the capacity to discover meaning through form, to organize attention, to perceive qualities and relationships before they become propositions, and to make judgments whose evidence remains substantially visual. This inquiry connects INsVI’s 2016 discussions of Kant and aesthetic judgment to the contemporary problem of building AI systems that do more than recognize and describe—that can learn how visual thought operates.

Contemporary art installation documented at Independent New York in 2016
2012–2022

New York Art Fair Research Circuit

Frieze New York · The Armory Show · Independent New York

INsVI maintained a field-observation practice across Frieze New York, The Armory Show, and Independent New York. These visits provided a recurring view of how contemporary work was installed, sequenced, mediated, and encountered outside museums: the density of fair architecture, the visual competition among booths, and the shifting relations among objects, audiences, galleries, and spectacle.

Visitor at Frieze Art Fair New York in 2013
Frieze New York, 2013. Global Good Group, CC BY 2.0, via Wikimedia Commons.
The Armory Show in New York in 2021
The Armory Show, 2021. Little doe run, CC BY-SA 4.0, via Wikimedia Commons.

Observed together over time, these fairs functioned as a comparative laboratory for visual culture. They revealed how scale, lighting, adjacency, circulation, branding, and attention affect judgment—and why visual intelligence must account for situated experience rather than treating an artwork as an isolated image file.

Archive note: The lead photograph documents Independent New York in 2016 from INsVI’s archive. Additional year-specific records will be added only when their dates and provenance are verified.

2017

Second Symposium

March 4 · Columbia University

INsVI’s second symposium, Cognitive Understanding of Visual Intelligence, considered how perception, cognition, and interpretation shape visual understanding.

Held from 2–6 PM in Room 602 of Columbia University’s Northwest Corner Building, the program brought philosophy into conversation with psychology and cognitive science. Keynote speakers were Dr. John Morrison, Assistant Professor of Philosophy at Barnard College and Columbia University, and Dr. Alexander Todorov, Professor of Psychology at Princeton University.

The program extended the inaugural symposium’s philosophical framework toward questions of cognitive architecture: how visual information becomes judgment, how faces and appearances are interpreted, and how forms of visual intelligence differ from linguistic reasoning.

2016

Inaugural Symposium

November 19 · Columbia University

INsVI’s inaugural symposium, Philosophical Understanding of Visual Intelligence, established visual intelligence as a question spanning aesthetics, philosophy, art, and artificial intelligence.

Held from 10:30 AM–5:30 PM in Room 633 of Columbia University’s Seeley W. Mudd Building, the symposium featured Prof. Ahmed Elgammal, Director of the Art and Artificial Intelligence Laboratory at Rutgers University; Dr. Gary Hatfield, Director of the Visual Studies Program at the University of Pennsylvania; and Dr. Sun-Joo Shin, Professor of Philosophy at Yale University. Dr. Elliot Paul of Barnard College and Columbia University served as faculty host.

The program placed machine vision beside longstanding philosophical problems: what it means to see, how visual experience becomes knowledge, whether aesthetic judgment can be formalized, and whether different kinds of intelligence apprehend images in fundamentally different ways.

Computational creativity scores for classical paintings from the Artchive dataset
2016

Foundational Conversations at Rutgers

July · Rutgers University

INsVI’s public record begins with an extended discussion with Prof. Ahmed Elgammal at Rutgers University. Elgammal is the founder and director of the Art and Artificial Intelligence Laboratory at Rutgers, whose stated aim was to advance artificial intelligence in the digital humanities by investigating perceptual and cognitive tasks related to human creativity.

The laboratory was working with large collections of digitized museum and art-historical images. Its research asked whether a machine could classify painting style, genre, and artist; learn a visual-similarity metric for fine-art paintings; trace possible paths of artistic influence; and assess creativity in relation to originality and influence over time. One study evaluated more than 62,000 paintings, while its artistic-influence dataset contained 1,710 high-resolution images by 66 artists across 13 styles and the years 1412–1996.

Against that background, the July 2016 exchange centered on Immanuel Kant’s account of aesthetic judgment and the possibility of different forms of visual intelligence. The discussion connected taste, imagination, perception, and art-historical interpretation with computational methods. It raised a question that remains central to INsVI: whether visual intelligence can be reduced to recognition and classification, or whether aesthetic judgment demands forms of cognition that current models do not capture.

Research context: Art and Artificial Intelligence Laboratory at Rutgers · Ahmed Elgammal, Rutgers University