RSS Amplifier

The RBDSai Lab’s Substack · May 29, 2026

Multimodal — How AI Holds Text and Image Together

0
Sign in to vote or save

Sahil Tanveer · The RBDSai Lab’s Substack

The popular shorthand for multimodal is that the AI can now see images. Not wrong — only shallow enough to make the next two years of architectural AI work harder than it needs to be. The deeper version: the model is aligning language, image, and spatial concept inside a shared meaning-space, and your prompt is a navigation instruction inside that space.

This is article four of a seven-part series. Article 2 introduced tokens, the unit of language for AI; embeddings extend that idea, so every concept becomes a coordinate and similar concepts cluster. Article 3 opened the diffusion engine; many image-generating systems use diffusion or diffusion-adjacent methods underneath, and what is new here is the joint coordinate system that lets text and image steer the same act of reasoning or generation. Once an architect can see it, “make it like this reference” stops being a prayer and starts behaving like a directable instrument.

The grammar is small. Vectors. Distance. Joint space. Dense regions and sparse ones.

The first correction is structural.

Inside the model, the word courtyard is not stored as letters. It is stored as a long list of numbers — typically several hundred to a few thousand, in a specific order. That list is called an embedding. The numbers are the coordinates of the word in a high-dimensional space.

You do not need to picture the dimensions. You need the intuition: every word, sentence, image, or image-region the model processes can be represented as a point in this space, and points with related meaning sit near each other. Courtyard lives close to atrium, light well, entry court. It lives far from parking deck, loading dock, service yard. Distance in the space corresponds, with surprising consistency, to closeness of meaning. [1]

Meaning has coordinates.

Engineers measure this distance with cosine similarity. You do not need the math. You need the picture: the embedding space is a navigable territory in which related concepts cluster and unrelated ones separate. It is the topography of meaning the model has learned from very large amounts of human material.

In our workshops, the architectural reader verifies this in five minutes. TensorFlow’s Embedding Projector lets you wander a high-dimensional space in 3D.[2] Words you know are related stand near each other without anyone having told the system they should.

Meaning is a coordinate. That single shift is the load-bearing primitive for everything that follows.

Embeddings on their own are not yet multimodal. A language-only model has its own coordinate system for words. An image-only model has its own coordinate system for images. The two systems do not speak.

The breakthrough was to teach them to.

In 2021, OpenAI published CLIPContrastive Language–Image Pretraining.[3] The setup is simple to describe and architecturally consequential. Take 400 million pairs of images and captions collected from the internet. Run each image through an image encoder; run each caption through a text encoder. Then train both encoders together with one rule: when a caption belongs to its image, push the two vectors together until they sit at almost the same coordinate. When a caption does not belong, push them apart.

After training, the result is a single space in which a photograph of a courtyard and the phrase “a photograph of a courtyard” land at roughly the same address. Image and text now share coordinates of meaning.[3]

This is the joint embedding idea, the quiet structural innovation that echoes through modern multimodal systems — Nano Banana Pro, gpt-image-2, Gemini’s image reasoning, Claude’s vision, the IP-Adapters layered on top of FLUX and Stable Diffusion. Different vendors, different architectures, same broad move: align text and image inside a shared territory so that anything you can locate with words can begin to be located with pictures.

Two encoders, one shared space. Trained to agree.

When you feed a multimodal model a photograph and a sentence, it is not “looking at the photo and reading the sentence.” It is computing two coordinates inside the same space and reasoning about the region they indicate.

The second mechanism builds on the first. Once text and image live in shared coordinates, the same attention machinery from Article 2—the junior architect glancing back across the brief—can attend to both at once.

In our reading, a modern multimodal model does not simply have a “language brain” and an “image brain” with a translator between. Increasingly, the direction is one window, one shared representation, one stream of attention. The image is converted into image tokens — patches or regions of the picture, projected into a space compatible with the language tokens. The text becomes language tokens. They are brought into the same reasoning window. Attention runs across the whole stream. [3]

That is what multimodal reasoning begins to become. Not merely a translation between two systems. A shared act of attention across compatible representations.

The implication for the architect is direct. When you upload a precedent image with your brief, the model is not “seeing the precedent and incorporating it.” It is placing the precedent’s coordinates inside the same window as the brief’s and letting attention find the region that satisfies both. This is one reason “make it like this reference” works. You are telling the model to start near that image’s coordinates and let your text tokens steer from there.

The space is not flat. That is the part most architects discover the hard way.

The training data the joint space was built from was not evenly distributed across what designers care about. Some regions were sampled enormously, the most Instagrammed buildings, the major architecture canon, anything heavily indexed by Pinterest and stock photography. Other regions were sampled thinly, vernacular Karnataka housing, Omani souk typologies, regional craft details, climatically specific monsoon strategies. The model has dense, well-mapped territory in some areas and sparse, unreliable territory in others. I think of the difference as dense versus sparse training territory, and it explains more daily friction than any single mechanism inside the model itself.

A prompt landing in dense territory behaves like a region with high-resolution streets. The model can be steered with subtle moves. “In the manner of a highly visible chiaroscuro clay textural language” lands in denser territory when the visual grammar is publicly abundant, consistently labelled, and easy for the model to triangulate. The model has a clearer neighbourhood. The architect can navigate it.

A prompt landing in sparse territory behaves like a region with no streets at all. The model interpolates. It guesses. It produces a confident-looking output drifting halfway toward whatever happens to be nearby. “In the manner of a 1972 KSCA-published Karnataka residential code-detail” lands in sparse territory because that material is barely represented. You will get something. It will not be what you asked for.

There is a second asymmetry, narrower but operationally important. Text-as-meaning lives in dense regions because the model has read enormous amounts of language. Text-as-image—rendering accurate labels onto a generated drawing—has historically lived in much sparser regions, because high-quality text-rendered-into-images was rarer in the training data. In our reading of the current data, that gap has narrowed sharply: Nano Banana Pro and gpt-image-2 render in-image text at a quality the previous generation could not approach. The architect’s annotated elevation is no longer impossible — but it remains a stress test current models still routinely fail at scale, especially with non-Latin scripts, dense labels, dimension strings, and the precise alignment of title-block fields.

The point is not the specific gap. The joint space is uneven, and the architect’s job is to learn which prompts navigate to dense regions and which ask the model to operate where it has no real coverage.

A worked example, because the abstraction will not land alone.

An architect uploads a single Zaha Hadid reference image — fluid geometry, continuous surfaces, compressed movement, structural drama — and pairs it with a brief: “This formal language, but for a four-bedroom Bengaluru courtyard house on a sixty-by-forty plot, climate-responsive, monsoon-aware, vernacular cues respected.”

What the multimodal system does next is not what the prompt-engineering vocabulary suggests.

It embeds the Zaha image. The image lands in the Zaha-adjacent neighbourhood the model can recognise because her architectural language is globally published, heavily indexed, visually distinctive, and consistently described across the internet. It embeds the brief. The brief lands in a different neighbourhood — courtyard-house typology, South Indian residential, climate-responsive, monsoon. Attention searches for the region that satisfies both anchors — Zaha-formal and Bengaluru-typological and monsoon-aware. From that region, generation samples.

If the architect has chosen well — a reference well-represented in public visual culture, a brief whose coordinates are reasonably mapped — the result is a courtyard house that carries Zaha’s formal register without losing the typological intelligence of the brief. If the architect has chosen badly — a niche reference, a brief pointing at sparse territory — the result is a confident-looking image drifting halfway toward whatever happens to be nearby. Same prompt, same model, two different outcomes, because the embeddings landed in two different regions of the space.

Same prompt. Two references. Two regions of the embedding space.

Same prompt. Two references. Two regions of embedding space. Different output.

That is the practical discipline. References should be chosen, in part, for their embedding signal—visually distinctive, publicly abundant, consistently labelled, and clearly addressable inside the joint space—not only for personal taste. A reference the model has barely seen produces noise. A reference the model has seen often, and seen consistently labelled, produces signal.

In studios we’ve worked with, the literate architect learns which is which. The slot-machine architect re-rolls.

Once the embedding-space picture is in place, several daily moves stop feeling like superstition.

“Make it like this reference” is no longer mystical. You are setting an embedding anchor. The model navigates from there.

“In the style of [obscure regional vernacular]” producing noise is not always a failure of prompt engineering. It is the model honestly reporting that the region you pointed at is sparsely covered. The fix is not a longer prompt; it is a denser anchor.

“Generate the section drawing with annotated dimensions” is not a guaranteed win even on the best 2026 models. Treat the output as a presentation-grade artefact that requires verification, not as a construction document.

The architect’s job is to operate the space, not to plead with the model. Prompt is a navigation instruction. Reference image is an embedding anchor. Iteration is a walk through the neighbourhood. The walk is more useful when you know what neighbourhood you are in.

The architect fluent in dense-versus-sparse will spot the openings as the next generation of models matures unevenly across architectural concerns. The architect who is not will read both surprises as luck.

It is not luck. It is the topography.

For how a studio can train its own version of this embedding space — making the model “know” its visual signature — read Part 5.

The embedding-space picture is teachable, but not by reading alone. It gets installed against your own live projects, where reference selection and the dense-versus-sparse instinct get rebuilt under deck pressure. We run a one-day workshop—the AI Fundamentals for Architects & Interior Designers workshop at Logika by RBDS AI Lab—for exactly this: the room where multimodal stops feeling magical and starts becoming a controllable instrument.

Explore Workshop

Sources

[1] Pinecone. Multi-modal ML with OpenAI’s CLIP — visual explainer of joint embeddings. https://www.pinecone.io/learn/series/image-search/clip/

Companion: Sebastian Raschka, Machine Learning Q and AI, Chapter 1. https://sebastianraschka.com/books/ml-q-and-ai-chapters/ch01/

[2] TensorFlow Embedding Projector — in-browser visualisation of high-dimensional embedding spaces. https://projector.tensorflow.org/

[3] Radford, Alec, et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP). OpenAI, 2021. https://openai.com/index/clip/ — Preprint: https://arxiv.org/abs/2103.00020 — GitHub: https://github.com/openai/CLIP

[4] Tanveer, Sahil. The Architecture of Memory: Navigating the Topography of Nano Banana 2. RBDS AI Lab Substack, March 3, 2026. https://rbdsailab.substack.com/p/the-architecture-of-memory-navigating

I’m Sahil Tanveer of the RBDS AI Lab, where we explore the evolving intersection of AI and Architecture through design practice, research, and public dialogue. If today’s post sparked your curiosity, here’s where you can dive deeper:

  • Read my bookDelirious Architecture: Midjourney for Architects, a 330-page exploration of AI’s role in design → Get it here

  • Join the conversation – Our WhatsApp Channel AI in Architecture shares mind-bending updates on AI’s impact on design → Follow here

  • Learn with us – Our online course AI Fundamentals for Lighting Designers is power-packed with 17+hrs of video content through 17 lessons → Enroll Here

  • Explore free resources – Setup guides, tools, and experiments on our Gumroad

  • Watch & listen – Our YouTube channel blends education with architectural art

  • Discover RBDS AI LabVisit our website

  • Speaking & events – I speak at conferences and universities across India and beyond. Past talks here

📩 Enquiries: sahil@rbdsailab.com | Instagram

Read the original on rbdsailab.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.