When the Kinect first landed in the studio at Art Zoyd in early February of 2011, it was still a curiosity. Microsoft had just released it for the Xbox, and while gamers were waving their arms in front of their televisions, I wasn’t a gamer. Never played video games (unless you count pinball). I wanted to point it at musicians. I picked one up used for cheap, and something I still do today: wait until people get bored with the latest gadget, then buy it second-hand.
The name itself, Kinect, was a little misleading. "Touch" was nowhere in the picture. It wasn’t about touching at all, it was about not touching, about reading the body and creating a dynamic map made of the joints of a skeleton turned into data from an image. At first I thought it was just a fancy video camera. Only later did I realize just how sophisticated it was, and even more amazing to think I could have all that for a few euros.
That was what interested me: the possibility of sculpting sound not with buttons or faders, but with the body itself. This interest went back almost fifteen years to my time at IRCAM. Despite all the technology available there, capturing movement was still more or less reserved for systems used in training military pilots: gyroscopes, early (and expensive) VR rigs, motion labs. At IRCAM I found workarounds, and some things came close, like Big Eye, an early motion-tracking system from STEIM (later absorbed into the far more sophisticated Isadora by my former colleague Mark Coniglio). Compared to what I could now do with the Kinect, Big Eye was basically the stone age. Even Max’s jit.cv library didn’t come close.
Having the device was one thing. Getting it to work was another. It was not plug-and-play, on the Mac you had to install stuff with names like OpenNI, the NiTE middleware, and sometimes wrappers like Synapse for Kinect and OSCeleton just to coax skeleton data into Max. That meant long downloads, installs, reboots, re-reboots, reinstalls… the usual ritual.
But once it worked, the magic happened. I built a simple ball-joint figure in LCD, each joint represented by a small sphere, so I could see the skeleton move in real time. As proof of concept, my first idea was obvious: a virtual Theremin — or two! Because the Kinect could capture multiple people at once, it could provide information on several musicians, or “bodies,” simultaneously.
So there we were: my colleague André Serre-Milan and I, both composers, flailing our arms in front of the Kinect, “playing” the theremin patch I’d built in Max/MSP. The patch showed a wonky multicolored ball-joint figure, tracking our hands and translating them into pitch and volume. Just for fun, I also tested the Z-axis, to see how well it tracked depth. It worked! I was so delighted I made a video and put it on YouTube to prove it.
Early Kinect experiments: virtual Theremin playing in Art Zoyd Studios, 2011
It was ridiculous (then and probably more so now), but we were young, adventurous, and a bit silly. Yet it was also the first glimpse of something: a limpid and natural way of letting movement conduct sound. I was so excited I even presented a paper on using it for a project with the ondes Martenot. The idea eventually fell away from that particular piece, but the Kinect stayed with me — in the studio, in my teaching at Brunel (where I often brought in game controllers as music interfaces), and even in seminars on music therapy.
Of course, the Kinect was not a cure-all. It suffered the same issues as any sensor: jitter, dropouts, false readings, data reduction. Frustrating, time-consuming — but possible, in ways I had only imagined before.
Later, I collaborated with Laurent Mariusse, a karate black belt and gifted percussionist. He wanted to perform a kata, Gankaku, with its sharp, precise, dynamic movements, using them as control gestures, a choreography that triggered and shaped sound in real time. Because he was also a composer and performer, we built a piece where he combined percussion (augmented with simple sensors) and movement in front of the Kinect.
We premiered the work in 2013 at the Festival CITYSONIC Mons. This was no longer silly — it was serious play. And it was fun.
Gangaku ou “le heron sur le rocher”, Festival Citysonic Mons 2013
Looking back now, there's an unsettling irony in how naturally we embraced a tool designed for surveillance. The same depth-sensing technology that let us sculpt sound with our bodies had been developed for very different purposes.
The Invisible Origins
As I was experimenting with these playful interactions in the studio, I had no idea where this technology actually came from. The Kinect hadn't simply emerged from Microsoft's gaming division, its story reached much deeper, back to Israeli military intelligence research that I would only learn about years later.
The technology was developed by PrimeSense, a company founded in 2005 by engineers with backgrounds in Israel's Unit 8200 and other R&D units. Originally, these systems were designed for reconnaissance, target acquisition, even “see-through-walls” surveillance. Military depth sensors could measure tens of meters with millimeter precision. The consumer Kinect was scaled down to about 3–4 meters… just enough for a living room.
Microsoft licensed PrimeSense's technology, acquired another Israeli company (3DV Systems), and invested half a billion dollars in what it branded Project Natal. The result: Kinect, the fastest-selling consumer electronic device in history, with ten million units sold in less than five months. Living rooms across the world were filled with people waving, jumping, dancing, completely unaware of the technology's military ancestry.
For me, though, the Kinect was never about gaming. It was about finding another way to perform with computers—not sitting behind a laptop, but moving through space, letting sound respond to the body. I had no sense then of the deeper implications.
The tension, of course, is that the same capacities that allow for expressive art — motion capture, gesture recognition, biometric sensing, are also the very ones used for control.
When we pick up tools like the Kinect in the studio, we inherit not only a piece of hardware, but a whole history: from military research labs, to consumer living rooms, and back again to surveillance systems.
Part 3 (coming soon) will turn to that broader picture: the implications of these technologies for users, for privacy, and for art. Because once you know that the skeleton dancing on your laptop is a cousin of the skeletons tracked in Xinjiang or London, you cannot unsee it.
And meanwhile, digging out my original Kinect, I find myself wondering again about its possibilities and how far Max (and my programming) have evolved since, and how tools like TouchDesigner have opened new avenues. I’m even eyeing the Kinect One, still cheap on the used market, but with time-of-flight depth sensing that takes the whole game up another level.
Background and Sources
This essay draws on research compiled from multiple sources documenting the Kinect's military origins, including reporting from Phys.org, Electronic Design, and TechCrunch, plus analytical work using AI tools to trace dual-use technology patterns. The fuller picture of surveillance applications will be explored in Part 3.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.