Studio Matrx Monthly · Volume 1 · Issue 4 · September 2026
Amogh N P
 In loving memory of Amogh N P — Architect · Designer · Visionary 
Spatial Input: Hands, Eyes, VoiceLesson 2.3
Spatial Computing for Design/Module 2 · The Technology

Lesson 2.3 · The Technology

Spatial Input: Hands, Eyes, Voice

How you act inside a three-dimensional world when there is no mouse and no menu bar - with controllers, bare hands and gesture, eye tracking and gaze, and voice - and the honest strengths and awkwardness of each

12 min Interactive lessonFree · open lessonByAmogh N P· Architect & interior designer
The hook

In a computed room there is no mouse to push and no menu to click - so how do you actually do anything? You point, pinch, look and speak. The mouse-and-menu era gives way to using your whole body, and each new channel is powerful and slightly awkward.

For forty years, interacting with a computer has meant one small, precise, learned motion: sliding a mouse across a desk to move an arrow, and clicking on words and icons arranged in menus. It is abstract - your hand is here, the cursor is over there - but it is fast, accurate and utterly familiar. Step inside a headset and that entire apparatus vanishes. There is no desk, no arrow, no menu bar pinned to the top of a rectangle. The digital content is all around you at arm's length and body scale. So the fundamental question of spatial computing is suddenly wide open again: when the world is three-dimensional and surrounds you, how do you actually reach into it and *do* something?

The answer is that input becomes embodied - you act with the parts of your body you already use to act on the real world. You hold controllers and press triggers; you reach out and pinch or grab with your bare hands; you simply *look* at a thing to point at it; you speak a command aloud. Each of these channels - controllers, hands, eyes, voice - opens a genuinely new way to work, and each carries its own strengths and its own awkwardness. This lesson takes them one at a time, honestly, and then steps back to the bigger shift they represent: from the precise, abstract mouse-and-menu to the natural, expressive, and still-imperfect business of interacting with your whole self.

Four channels: controllers (precise) | hands+gesture (immediate) | gaze (effortless targeting) | voice (hands-free naming). Symbolic -> embodied. A different bargain, not an upgrade. Combine them.

Controllers - the precise, learnable starting point

The most established way to act in a computed world is a pair of hand-held controllers, one in each hand, tracked in full 6DoF so their virtual position matches your real hand exactly. They carry buttons, triggers, thumbsticks and touch-sensitive pads, and they usually offer a small burst of vibration - simple haptic feedback - when you press, grab or bump something, giving a satisfying physical confirmation that your action registered. Point a controller and a laser-like ray extends from it into the scene; aim the ray at a distant object or a floating button and pull the trigger to select or grab it. Reach out and the controller becomes your hand for picking up and moving things nearby.

Controllers earn their place through precision and reliability. A physical trigger has an unambiguous pressed-or-not state, so selection is crisp and error-free in a way that reading intent from a moving hand is not. The pointing ray is accurate over distance. The buttons give you a rich vocabulary of distinct actions without guesswork, and the haptic buzz closes the loop so you feel that something happened. For any task that needs confident, repeatable selection and manipulation - grabbing a wall and dragging it, clicking through a review checklist, drawing a precise line in space - controllers remain the dependable workhorse, which is why they have not disappeared even as hand tracking has improved.

The costs are equally honest. Controllers are extra objects you must pick up, hold and not put down - your hands are occupied, so you cannot simultaneously hold a pen, a phone or a real drawing. They have batteries that die. They must be learned: which button does what is not obvious to a first-time user, so a client handed controllers often fumbles before they can do anything, which undercuts the very immediacy that makes immersion valuable for communication. And they insert a small layer of abstraction back between you and the world - you are pressing a device, not touching the thing. For a designer running a client session, that learning curve is the key limitation; for your own repeated, precise work, the precision usually wins. Controllers are the solid default you reach for when accuracy matters more than immediacy.

FOUR CHANNELS OF SPATIAL INPUTCONTROLLERS+ precise, reliable, buttons + haptic- hands occupied, batteries- must be learned (clients fumble)HANDS + GESTURE+ immediate, nothing to learn- flickers, guesses intent- no haptics, tiring armsGAZE (eye tracking)+ effortless targeting, foveated render- restless (Midas touch)- privacy considerationsVOICE+ name it hands-free, skip menus- noise, accents, imprecise- socially odd, privacy
Zoom
The four channels of spatial input at a glance: controllers, bare-hand gesture, gaze and voice, each with its core strength and its core awkwardness. No channel is complete on its own, which is why strong interfaces combine them.

Controllers = buttons + trigger + a pointing ray + a haptic buzz. Precise and reliable, but occupy your hands, need batteries, and must be learned (clients fumble).

Hand tracking and gesture - reaching in with nothing in your hands

The most natural-feeling way to act is to use no device at all. Hand tracking uses the headset's outward cameras to see your bare hands - the position and bend of every finger joint - and reconstruct them live inside the computed world, so you look down and see your own hands, and reach out and touch, push or grab virtual things directly. Layered on top is gesture recognition: the software learns to recognise specific hand shapes and movements as commands. The near-universal one is the pinch - touching thumb to forefinger - which acts as the spatial equivalent of a click: look or point at something, pinch to select it, pinch and move to drag it. Others include grabbing a fist to pick up, spreading the hands to scale, or an open palm to summon a menu.

The appeal is immediacy and accessibility. Nothing to pick up, nothing to learn in the way controllers must be learned - a first-time user, a client, a child instinctively reaches out and touches, which is exactly the frictionless directness that makes immersion so persuasive for communication. Your hands are free between actions to gesture, point at a colleague, or hold a real object. For handing an experience to a non-specialist, bare hands remove a real barrier.

But the honesty of this section matters most here, because hand tracking is where hype most outruns reality. Reading precise intent from a moving hand is genuinely hard: the cameras can lose a hand that moves out of view, is turned edge-on, or is poorly lit, so it flickers or freezes. There is no physical button, so the system must *guess* when a pinch is deliberate versus incidental, which produces both missed actions and accidental ones. Crucially, there is no haptic feedback - your fingers close on empty air, so you never feel the object you are 'touching', which makes precise manipulation oddly effortful and tiring, and holding your arms up to work in front of you brings on 'gorilla arm' fatigue quickly. Hand tracking is wonderful for casual, expressive, low-precision interaction and for first-contact accessibility, and frustrating for sustained precise work - which is exactly why serious use often mixes it with controllers rather than replacing them.

HAND GESTURES: shapes become commandsPINCH = clickthumb to forefingerGRAB = pick upclose a fistSPREAD = scalehands apartNo haptic feedback: fingers close on empty air. Hands out of view or poorly lit flicker or freeze.But: nothing to pick up, nothing to learn - a client just reaches out and touches.
Zoom
Bare-hand tracking and gesture: the headset cameras reconstruct your hands, and specific shapes become commands - the pinch as a click, a grab to pick up, spreading the hands to scale. Immediate and learnable-free, but with no haptic feedback your fingers close on empty air, and unseen or poorly lit hands flicker.

Bare hands seen by the cameras + gestures (pinch = click, grab, spread to scale). Immediate, nothing to learn - but flickers when unseen, guesses intent, no haptics (touching air), tiring arms.

Eye tracking and gaze, and voice - looking and speaking as input

Two more channels turn parts of you that are not even hands into input. Eye tracking uses tiny inward-facing cameras to follow where your eyes are pointed, so the system knows what you are *looking at* moment to moment. As an interaction this is uncannily fast and effortless: you look at a button and it highlights; you look at an object and pinch, and the look did the pointing while the pinch did the clicking - a 'look-and-pinch' pairing that feels almost like the interface reading your mind, because targeting by gaze is faster than moving any hand or controller. Eye tracking also quietly powers foveated rendering, where the device draws only the small patch you are actually looking at in full detail and spends less effort on the blurry periphery, saving precious computing power. Its awkwardness is subtler: gaze is restless and not always intentional - you look at things you are not choosing - so using a look as a deliberate command risks the 'Midas touch' problem, where everything you glance at reacts. And because it senses where you look, it raises real privacy considerations that stay with manufacturers' guidance.

Voice is the fourth channel: simply speaking a command aloud. Its strength is that it bypasses all the pointing entirely - naming a thing or an action ('show the second floor', 'make this wall red', 'measure this') can be far faster than navigating to it by hand, especially for summoning tools buried in menus, and it leaves your hands and gaze free for other work. It is also accessible, needing no learned button vocabulary. Its awkwardness is equally plain: it is unreliable in noisy rooms and across accents (a real concern for Indian English and multilingual practices), it feels socially odd to talk aloud to a device in a meeting, it struggles with precise or spatial instructions ('move it a bit that way'), and it raises the same privacy questions as any always-listening microphone.

The honest conclusion for both is that neither is a complete input method on its own. Gaze and voice shine as *fast complements* - gaze for effortless targeting, voice for naming and summoning - woven together with a hand or controller that supplies the deliberate, precise confirming action. The best spatial interfaces do not pick one channel; they combine the strengths of several while covering each one's awkwardness.

GAZE + VOICE: fast complements, confirmed by a handGAZE (look to target)targetlook targets + pinch confirms+ effortless, powers foveated render- restless: the Midas touchVOICE (name to do)"show the second floor""make this wall red"+ hands free, skips menus- noise, accents, imprecise- socially odd, privacyNeither is complete alone - weave gaze + voice with a confirming hand or controller.Privacy of gaze and voice sensing stays with the manufacturers' guidance.
Zoom
Gaze and voice as fast complements: eye tracking targets what you look at (and powers foveated rendering) while voice names an action hands-free - each paired with a deliberate pinch or trigger to confirm. Neither is complete alone: gaze is restless (the Midas touch) and voice stumbles on noise, accents and precise spatial instructions.

Eyes: look to target (fast, effortless) + powers foveated rendering; but restless gaze = Midas-touch + privacy. Voice: name it to do it, hands free; but noise/accents, socially odd, imprecise, privacy. Both = complements, not complete.

The shift to embodied interaction - and why no single channel wins

Step back and the deeper change is not any one of these channels but what they add up to: a move from symbolic interaction to embodied interaction. The mouse-and-menu is symbolic - you manipulate an abstract cursor and pick words from lists, a learned code standing in for action, brilliant for precise, discrete tasks on a flat screen. Spatial input is embodied - you reach, grab, look and speak, using the same faculties you use to act on the real world, so the interface can feel astonishingly natural and requires far less translation in your head. This is the promise captured by the term natural user interface: interaction that draws on skills you were born with rather than ones you had to learn. For letting a non-designer simply *do* something in a model, that naturalness is a genuine gift.

But 'natural' is a claim to weigh, not to swallow. Embodied input is expressive and immediate, yet it is often *less* precise, *less* reliable and *more tiring* than the humble mouse for real work. A mouse never suffers gorilla arm, never loses tracking, never guesses your intent, never mishears you, and gives pixel-accurate control for hours - which is exactly why, for detailed design, drawing and documentation, the flat screen and its abstract, precise input remain faster and better, as the whole course keeps insisting. Embodied interaction is not a strict upgrade; it is a different bargain, trading precision and stamina for immediacy and presence.

So the practical wisdom mirrors the device lesson. No single channel wins; the strong spatial interfaces combine them - gaze to target, a pinch or trigger to confirm, voice to summon, hands or controllers to manipulate - each covering another's weakness. As a designer you choose input the way you choose a device: by the task. Reach for bare hands and voice when immediacy and accessibility matter most (a client touching their space); reach for controllers when precision and reliability matter most (your own careful manipulation); and reach back for the mouse and screen when the task is detailed, sustained work that embodied input simply does less well. And through all of it, the channel is only how you *act on* the model - the model itself remains a tool for seeing and communicating, never the source of truth for binding results, which stay with the verified drawings, survey and specialists.

SYMBOLIC -> EMBODIED: a different bargain, not a strict upgradeSYMBOLIC (mouse + menu)abstract cursor, lists of words+ pixel-accurate, reliable+ low fatigue, hours of work- learned code, not naturalstill wins for drawing, detailing,documentation->EMBODIED (reach/look/speak)natural user interface+ immediate, expressive+ accessible - anyone can act- less precise, less reliable- more tiringstrong interfaces combine channels
Zoom
The deeper shift: from symbolic mouse-and-menu (an abstract cursor and lists, precise and low-fatigue) to embodied spatial input (reach, look, speak, natural and immediate). It is a different bargain - trading precision and stamina for immediacy and presence - not a strict upgrade, which is why the screen still wins for detailed work.
Verify-this: match the input channel to the task, and combine channels to cover each one's weakness

Controllers

Precise, reliable manipulation

6DoF hand-held devices with buttons, trigger, pointing ray and haptic feedback. Crisp, dependable selection - the workhorse for precise, repeated work - but occupy the hands, need batteries and must be learned (clients fumble).

Hand tracking + gesture

Immediate, learnable-free action

Cameras see bare hands; pinch = click, grab, spread to scale. Superb for accessibility and first contact; but flickers when unseen, guesses intent, gives no haptics and tires the arms. Great for casual, poor for sustained precision.

Gaze + voice

Fast complements, not complete methods

Eye tracking targets effortlessly (and powers foveated rendering) but risks the Midas touch; voice names and summons hands-free but stumbles on noise, accents (Indian English) and precise spatial instructions. Best woven with a confirming hand or controller. Privacy stays with manufacturers.

Embodied, not a source of truth

What the input decides

Spatial input is how you act on the model - a different bargain from the mouse (immediacy for precision), not a strict upgrade. However you interact, the model stays a tool for seeing; binding dimensions and decisions stay with the verified drawings, survey and specialists.

Hands-on workshop

Workshop — map the input channels to a real interaction

Understanding spatial input means feeling why each channel fits some actions and fights others. This workshop takes a concrete design interaction and reasons through how you would perform it with each channel - exposing the strengths and awkwardness first-hand, and building the instinct to combine channels rather than force one.

A notebook and this lesson. If a hand-tracking headset is available, actually attempting the task with each channel makes the strengths and awkwardness vivid - but the reasoning works fully without one.

Given & goal
Goal: build judgement about which input channel fits which action, and why the strongest interfaces combine them
Inputs: this lesson + a notebook (a headset with hand tracking helps but is not required)
Time: ~35 minutes
  1. 1Pick one concrete immersive task with several sub-actions - for example, a client review of a room where you must: select a wall, change its finish, move a piece of furniture, summon a menu, and step to the next layout option.
  2. 2For each sub-action, write how you would do it with a controller, with bare-hand gesture, with gaze, and with voice - and rate each as good, awkward or bad, with a one-line reason.
  3. 3Notice the pattern: which sub-actions want precision (favouring controllers), which want immediacy and accessibility (favouring hands/voice), and which want effortless targeting (favouring gaze).
  4. 4Design a combined scheme: assign each sub-action to the channel (or pair of channels, e.g. gaze-to-target plus pinch-to-confirm) that fits it best, and note how the combination covers each channel's weakness.
  5. 5Write a short reflection on the honest limits you found - fatigue, flicker, guessed intent, the Midas touch, noise and accents - and confirm one binding result that stays off every channel and with the drawings and specialists.

You’ll walk away with
A one-page interaction map for a real immersive task: each sub-action matched to the best input channel(s) with a reason, an honest note on the awkwardness of each channel you encountered, and a clear statement that input acts on the model only while binding results stay with the verified drawings and specialists.

The worked example

Three altitudes on the same idea

Read the band that fits you — or all three.

For the architectDesigning, reviewing and communicating buildings in immersive 3D - where it earns its place

Choose input by the task, and expect to combine channels rather than crown one. For your own repeated, precise work in a model - grabbing and moving elements, drawing in space, stepping through a review - controllers give the reliable, accurate, button-rich control that bare hands still cannot match, and the haptic confirmation matters. For handing an experience to a client or consultant, bare-hand tracking removes the learning barrier so they can simply reach out and touch, and voice can summon tools without menu-hunting. Weave gaze (effortless targeting) and voice (naming and summoning) as fast complements to a confirming pinch or trigger. But keep the honest ceiling in view: embodied input is immediate yet less precise, less reliable and more tiring than a mouse, so for detailed drawing, detailing and documentation the flat screen still wins - which is why immersion earns its place task by task. And whichever channel you use, it only acts on the model; binding dimensions and decisions stay with the verified drawings, survey and specialists.

For the interior designerLetting clients stand inside a space at true scale before it is built

For client sessions, bare hands and voice are usually your friends; controllers are for your own precise work. The magic of an immersive interior review is a client instinctively reaching out to touch a surface or move toward the window - hand tracking delivers that with nothing to learn, which is exactly why it suits first-time users who would fumble with controllers. Add simple voice ('show the darker floor', 'next layout') to switch options without breaking the spell, and let gaze do effortless targeting. Keep the interactions few and obvious - a pinch to select, a look to point - so a nervous client is never lost. For your own careful placement and adjustment work, controllers give the precision and reliable selection bare hands lack. Be honest that hand tracking can flicker and tire, that there is no real touch (fingers close on air), and that voice stumbles in noisy rooms and across accents. The input is how the client explores the space - the binding specification still lives with the drawings and samples.

For the studentHow the computer leaves the screen - and where XR genuinely helps design and where it does not

Learn the four channels and, more importantly, the shift they represent. Controllers (buttons, trigger, pointing ray, a haptic buzz) are precise and reliable but occupy your hands and must be learned. Hand tracking and gesture (pinch to click, grab, spread to scale) are immediate and need no learning, but flicker when unseen, guess your intent, give no haptic feedback and tire your arms. Eye tracking lets you target by simply looking (and quietly powers foveated rendering) but restless gaze risks the 'Midas touch'. Voice lets you name an action hands-free but stumbles on noise, accents and precise spatial instructions. The big idea underneath is the move from symbolic mouse-and-menu to embodied interaction - a natural user interface that uses faculties you were born with. Hold it honestly: embodied input is more immediate but often less precise and more tiring than a mouse, so it is a different bargain, not a strict upgrade - and the strongest interfaces combine channels, each covering another's weakness.

Misconception check

In spatial computing you just use your hands and voice naturally, the way you do in real life, so interacting in 3D is effortless and obviously better than the fiddly old mouse and keyboard - and controllers are just a clumsy leftover that hand tracking has made obsolete.

The 'natural, therefore effortless and better' story is half true and importantly misleading. Embodied input - reaching, pinching, looking, speaking - genuinely can feel more natural than an abstract cursor, and for immediacy, expressiveness and letting a non-specialist simply act, it is a real gift; that is the promise of a natural user interface. But 'natural' is not the same as 'precise, reliable or comfortable'. Bare-hand tracking loses your hand when it moves out of view or is poorly lit, has to guess whether a pinch is deliberate, gives no haptic feedback (your fingers close on empty air), and tires your arms quickly ('gorilla arm'). Gaze is fast but restless, risking the 'Midas touch' where everything you glance at reacts. Voice is hands-free but stumbles in noisy rooms and across accents, feels socially odd, and struggles with precise spatial instructions. Against all this, the humble mouse gives pixel-accurate, reliable, low-fatigue control for hours - which is exactly why controllers persist (their physical trigger and buttons are crisp and dependable where hands are not) and why the flat screen still wins for detailed drawing, detailing and documentation. Embodied interaction is a different bargain - trading precision and stamina for immediacy and presence - not a strict upgrade, and the strongest spatial interfaces combine several channels so each covers another's weakness. Match the input to the task, exactly as you match the device to the task.
Try it

Do it yourself

You can reason this out without a headset - but try a channel if you have one.

  1. 1Name the four spatial-input channels and give one strength and one awkwardness of each.
  2. 2Why do controllers persist despite good hand tracking? What does a physical trigger and haptic feedback give that bare hands do not?
  3. 3Explain the 'look-and-pinch' pairing and why combining gaze with a hand action beats using either alone.
  4. 4What is the difference between symbolic (mouse-and-menu) and embodied interaction, and why is embodied not a strict upgrade?
  5. 5Give a client-session action best done with hands or voice, and an own-work action best done with controllers - and say why.
Take this with you

The one line to carry out

Inside a computed world there is no mouse and no menu, so you act with your body across four channels - controllers (precise, reliable, but hands occupied and must be learned), bare-hand gesture (immediate and learnable-free, but flickery, intent-guessing, haptic-less and tiring), gaze (effortless targeting, but restless) and voice (hands-free naming, but noise-, accent- and precision-limited) - which is a move from symbolic to embodied, natural interaction that is more immediate but often less precise and more tiring than a mouse, so it is a different bargain, not a strict upgrade, and the strongest interfaces combine channels to cover each one's weakness.
Take it further
References & further reading

Peer-reviewed journals & authoritative standards

  1. 013D interactionWikipedia — 3D interaction, 2026.
  2. 02Gesture recognitionWikipedia — Gesture recognition, 2026.
  3. 03Eye trackingWikipedia — Eye tracking, 2026.
  4. 04Hand trackingWikipedia — Hand tracking, 2026.
  5. 05Natural user interfaceWikipedia — Natural user interface, 2026.
Related lessons
Recap
Inside a headset the mouse and the menu bar disappear, and interaction becomes embodied: you act with the parts of your body you already use on the real world, across four channels, each powerful and imperfect. Hand-held controllers, tracked in 6DoF with buttons, a trigger, a pointing ray and a haptic buzz, give precise, reliable, learnable selection and manipulation - the workhorse for careful work - but they occupy your hands, need batteries, and must be learned, so first-time users fumble. Bare-hand tracking with gesture (the pinch as a click, grab to pick up, spread to scale) is immediate and needs no learning, superb for accessibility and first contact, but it flickers when a hand is unseen or poorly lit, must guess your intent, gives no haptic feedback so your fingers close on empty air, and tires the arms. Eye tracking lets you target simply by looking - uncannily fast, and it quietly powers foveated rendering - but restless gaze risks the 'Midas touch', and it raises privacy questions. Voice lets you name an action hands-free and skip menu-hunting, but stumbles in noisy rooms and across accents like Indian English, feels socially odd, and struggles with precise spatial instructions. Underneath sits the real shift, from symbolic mouse-and-menu to embodied, natural interaction - a genuine gift for immediacy and accessibility, but honestly a different bargain, not a strict upgrade, since embodied input is often less precise and more tiring than the humble mouse, which is why the flat screen still wins for detailed work. So you choose input by the task, the strongest interfaces combine channels to cover each one's weakness, and however you interact, the model stays a tool for seeing while binding results stay with the drawings, survey and specialists.
Carry forward →

You now know how the illusion is built, what device families deliver it and how you act inside it. The honest counterweight comes next: the real, unavoidable limits of today's hardware - weight, heat, field of view, resolution, battery, comfort over time and the vergence-accommodation conflict - and why comfortable all-day hardware still does not exist. Next, the limits, and the discipline they demand.

A

The author

Amogh N P

Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.

More about Amogh →