Studio Matrx Monthly · Volume 1 · Issue 3 · August 2026
Amogh N P
 In loving memory of Amogh N P — Architect · Designer · Visionary 
Sketch-to-Image & ControlNetLesson 3.2
AID for Architecture, Planning & Urban Design/Module 3 · AI for Concept & Ideation

Lesson 3.2 · AI for Concept & Ideation

Sketch-to-Image & ControlNet

Turning your sketches and massing into rendered concepts with img2img and ControlNet - keeping YOUR geometry while the AI explores only the look

13 min Interactive lessonFree · open lessonByAmogh N P· Architect & interior designer
The hook

What if the AI kept the building you drew, and only changed how it looks?

Plain text-to-image is a slot machine: it invents a new form every run, and none of them is the one you carefully composed. That is fine for scattering early ideas - but useless the moment you have a massing you believe in and want to see it in different materials and light.

That is what img2img and ControlNet are for. They feed your drawing into the model as a constraint, so the AI keeps your geometry - your roofline, your window rhythm, your plan - and explores only the surface: material, mood, atmosphere. This is the lesson where AI rendering stops overwriting your design and starts serving it.

ControlNet = AI as a renderer of YOUR design, not a generator of its own. Keep the form; explore the skin.

Why constraint is the whole point

In the last lesson the AI decided the form and you reacted. That is perfect for opening the solution space, but the opposite of what you want once you have committed to a composition. When you have a massing model, a plan, or a considered perspective sketch, you do not want the AI to reinvent it - you want it to render it. The problem is that a text prompt alone gives the model no idea of your geometry, so it will cheerfully produce a different, prettier building every time.

Image-to-image (img2img) solves half of this. Instead of starting from pure noise, the model starts from your image and partially denoises it toward your prompt. A single dial - variously called denoise strength or image weight - controls how far it drifts. Low strength (say 0.2-0.35) barely changes your input, just restyling surfaces; high strength (0.7+) keeps only a loose echo of your composition and invents the rest. Learning to feel that dial is most of img2img.

The honest limit: img2img respects colour and rough layout, but under a strong prompt it will still bend edges, move openings and hallucinate detail. For tight control - where a roofline or a column grid must survive exactly - you need something that reads structure, not just pixels. That is ControlNet.

It is worth being clear about why you would want constraint at all, because it feels like giving up the AI's magic. The answer is ownership. A text-to-image concept, however beautiful, is a suggestion the model made; a constrained render is your design, seen in a new light. As a project matures, the second is almost always what you need - to show a client the scheme you agreed, to test materials on the massing you resolved, to keep a competition entry recognisably one building across ten images. Constraint is not a limitation on creativity; it is what lets you apply the AI's speed and beauty to a decision you have already earned, instead of endlessly re-opening it.

THE CONTROL SPECTRUMLOOSE / AI DECIDES FORMTIGHT / YOU DECIDE FORMtext-to-imagewords only;wild varietyimg2img (low)keeps roughlayout + colourimg2img (high)close to input;restyle onlyControlNetlocks edges /depth exactlyEarly = go loose for surprise. Later = go tight to protect the design you have committed to.Choose your point on the dial by how much of the form you have already decided.
Zoom
The control spectrum from loose to tight: text-to-image lets the AI decide the form; img2img keeps more of your input as denoise strength drops; ControlNet locks edges or depth exactly. Choose your point on the dial by how much of the form you have already committed to.

img2img dial: low denoise = restyle only; high denoise = keep a loose echo and invent the rest.

ControlNet: locking the geometry you drew

ControlNet is an add-on to Stable Diffusion that conditions the image on a structural map extracted from your input, not just its colours. It comes in flavours, each reading a different kind of structure. Canny and Scribble follow edges and lines - ideal for a hand sketch or a line export from CAD. Depth reads near-versus-far - perfect for a grey massing model or a SketchUp view, so the AI understands the volumes. MLSD snaps to straight lines, which keeps architectural geometry crisp. Segmentation maps regions (wall, glass, sky, ground) so you can say what each area is.

The workflow is straightforward in ComfyUI or Automatic1111: load your sketch or massing render, pick a ControlNet type, and it produces a control map the model must honour while your text prompt supplies material, light and mood. Because the geometry is locked to your input, you can now run ten variations - warm timber at dusk, white render at noon, weathered concrete under cloud - and every one keeps your building.

text
Input:   grey SketchUp massing, eye-level view
Control: Depth (reads the volumes)
Prompt:  rammed-earth facade, deep window reveals, warm
         late-afternoon light, architectural photograph,
         no people, no text
Result:  your massing, now rendered - same form, new skin

The key mental shift: ControlNet makes the AI a renderer of your design, not a generator of its own. The composition is yours; the AI only dresses it.

A few practical notes save frustration. The quality of the control map matters: a clean line drawing or a clear grey massing with readable depth gives the model something firm to hold, while a muddy, cluttered input produces muddy, uncertain output. You can dial each ControlNet's strength or weight - how strictly the model must obey the map - so a lower weight lets the prompt breathe while a higher one clamps the geometry hard. And you can be selective about which flavour to trust: Depth is forgiving and great for volumes but vague about edges; Canny is precise about lines but ignores depth. Choosing the map that matches what you most need to protect - the silhouette, the volumes, the window grid - is half the craft of getting a faithful render out of your own sketch.

KEEP YOUR GEOMETRY, EXPLORE THE LOOKYOUR SKETCHCONTROLNETreads edges /depth / lineslocks the formRENDERED CONCEPTsame shape, new look+ prompt: material, light, moodControlNet constrains the AI to YOUR composition - it dresses the geometry, it does not redraw it.
Zoom
The sketch-to-image workflow: ControlNet reads the edges or depth of your sketch or massing and locks the form, while the prompt supplies material, light and mood. Same shape in, new look out - the AI dresses your geometry, it does not redraw it.

Choosing where to sit on the control dial

There is a spectrum from loose to tight, and the right point depends on how much of the form you have already decided. Pure text-to-image is loosest - the AI owns the form; use it to diverge when you have nothing committed. img2img at low strength keeps your rough layout and colour, good for quickly reskinning a concept. img2img at high strength hews close to your input. ControlNet is tightest, pinning edges or depth exactly - use it when the geometry is sacred.

A practical rule: the more of the design you have committed to, the tighter you should go. Early, go loose and let the model surprise you. Once you have a massing you believe in, go to ControlNet-Depth so exploration cannot quietly erode the decision you made. Sitting at the wrong point wastes effort - loose control on a committed design gives you pretty images of the wrong building; tight control too early kills the surprise you were hoping for.

You can also stack constraints. Combining a ControlNet map with a modest img2img strength, or two ControlNets (Depth for volume plus Canny for edges), gives fine-grained hold - though every added constraint narrows variety, so add them only as far as the design demands. As always, more scrutiny where the stakes are higher: a mood study can drift; a client-facing image of an agreed scheme should keep the scheme.

There is a natural progression across a project's life that maps onto this dial. In week one you are at the loose end, letting text-to-image throw forms at you. By schematic design you have a massing, so you slide toward img2img and light ControlNet to reskin it. By the time you are presenting a resolved scheme you are at the tight end, ControlNet-Depth locking the building while you explore only material and light. Reading where a project sits and picking the matching control level is a judgement in itself - and getting it wrong is one of the commonest ways designers waste hours with these tools, generating gorgeous images of a building nobody asked for.

THE CONTROL SPECTRUMLOOSE / AI DECIDES FORMTIGHT / YOU DECIDE FORMtext-to-imagewords only;wild varietyimg2img (low)keeps roughlayout + colourimg2img (high)close to input;restyle onlyControlNetlocks edges /depth exactlyEarly = go loose for surprise. Later = go tight to protect the design you have committed to.Choose your point on the dial by how much of the form you have already decided.
Zoom
The control spectrum from loose to tight: text-to-image lets the AI decide the form; img2img keeps more of your input as denoise strength drops; ControlNet locks edges or depth exactly. Choose your point on the dial by how much of the form you have already committed to.

More of the design decided = go tighter. Committed massing? ControlNet-Depth so exploration cannot erode it.

Where this fits your process - and where it stops

This technique sits in the fertile middle of a project: past pure ideation, before real visualization. It shines at taking a rough SketchUp or Rhino massing and making it legible and atmospheric enough to discuss - with a client, a team, a jury - without the cost of a full render. It is also a fast way to test material and light options against a fixed form: same geometry, ten skins, in the time a single V-Ray render would take.

But keep the boundaries honest. ControlNet holds your composition, not your dimensions - it will not keep a window at exactly 1200mm or guarantee a buildable detail, and it still invents texture and small features. The output is a persuasive concept image, not a construction document and not a substitute for a measured drawing. Where accuracy is contractual - a planning submission, a heritage context - a diffusion image is evidence of intent, not fact.

And the authorship principle from the last lesson still holds, now with a stronger foundation: because the geometry started as yours, the result is far more clearly your design than a text-to-image guess ever could be. That is the quiet reason this workflow matters - it lets you use the speed and beauty of diffusion while the building on the screen remains the one you drew. Module 5 picks this up for finished visualization; here it stays a concept tool.

One workflow habit ties it all together: keep the source of truth in your CAD or BIM model and treat every diffusion image as a disposable view off that model. When the design changes, you change the model and re-export the control input, rather than trying to edit the AI image. That discipline stops the common trap of a beautiful render quietly becoming the design of record - a picture nobody can dimension, coordinate, or build from. Used this way, sketch-to-image is a fast, honest window onto a design that lives properly elsewhere, and that is exactly the role it should keep.

Techniques you will meet in this lesson

img2img

Generating from your image plus a prompt, not from noise

Denoise strength dials how far it drifts from your input - low restyles, high reinvents. The gentlest form of control.

ControlNet

Conditioning Stable Diffusion on a structural map from your input

The core tool for keeping your geometry. Comes in types (Canny, Depth, MLSD, Segmentation) for different inputs.

ControlNet Depth

Reads near/far from a massing model or SketchUp view

Best for turning grey volumes into rendered concepts - the AI understands the 3D form, not just outlines.

Canny / Scribble

Edge- and line-following ControlNet types

Ideal for hand sketches and CAD line exports where the outline must be honoured exactly.

Denoise strength

The img2img dial for how much the output deviates

The single most important lever in sketch-to-image; feel it out on a fixed input before trusting it.

Hands-on workshop

Workshop — render a form you drew

You will take one piece of geometry you authored - a hand sketch, a CAD line drawing, or a grey massing render - and produce several rendered concepts that all keep that form, using img2img or ControlNet. The aim is to feel the control dial and prove the building on screen is still yours.

Stable Diffusion with the ControlNet extension (ComfyUI or Automatic1111 - free, local) is ideal; a hosted img2img tool works for the img2img steps. Plus your own sketch or massing view.

Given & goal
Goal: get multiple material/light options of ONE form you designed
Inputs: a sketch or grey massing image + Stable Diffusion with ControlNet (ComfyUI/Automatic1111), or an img2img tool
Time: ~45 minutes
  1. 1Prepare one input image: a clean line sketch, a CAD edge export, or a grey SketchUp/Rhino massing view (eye-level reads best).
  2. 2Start with img2img: set denoise low (~0.3) with a material+light prompt, then rerun at ~0.6. Note how much your form survives at each.
  3. 3Switch to ControlNet - Depth for a massing, Canny/Scribble for a sketch - and generate with the same prompt. Compare how much better it holds your geometry.
  4. 4Hold the ControlNet input fixed and change only the prompt three times (three materials or three times of day). Confirm every output keeps your form.
  5. 5Pick the strongest image, and in one sentence state what stayed yours (composition) and what the AI supplied (skin, light, mood).

You’ll walk away with
A small comparison sheet: your input, one img2img result at low and high denoise, and three ControlNet variations of the same form - with a note on which control level held your design best and why.

The worked example

Three altitudes on the same idea

Read the band that fits you — or all three.

For the architectAI across the whole design process

This is the bridge from your massing model to a conversation-ready image. Export a grey SketchUp or Rhino view, run ControlNet-Depth, and show a client the agreed scheme in three material and light options in an afternoon - no render farm. Use Canny or MLSD from a CAD line export when the elevation geometry must stay crisp. Just remember it holds composition, not dimensions: it is a concept renderer, not a drawing.

For the interior designerAI for ideation, specs & client work

Photograph or model the actual room, then restyle it against locked geometry. img2img at low denoise, or ControlNet from a perspective sketch, lets you show a client THEIR space in a new palette, finish or lighting mood while the walls, windows and proportions stay put. It is far more convincing than a generic concept because it is recognisably their room - a powerful, honest way to present material and mood options.

For the studentAn AI-fluent design skillset

ControlNet is how you make studio work look resolved without faking the design. Instead of a generic AI render that a tutor can tell you did not author, run your own massing or hand sketch through Depth or Scribble - the form is provably yours, the AI only renders it. Learn the denoise dial and the ControlNet types now; they are among the most employable diffusion skills, and they keep your voice in the image.

Misconception check

img2img and ControlNet let AI render my design accurately, so I can trust the geometry in the output.

They constrain the AI to your composition, which is a big step up from text-to-image - but they hold shape and layout, not measurements. ControlNet-Depth or Canny will keep your massing, roofline and window rhythm recognisable, yet the model still invents texture, bends fine detail, and honours no real dimension: a reveal will not be exactly 1200mm and an output is never a measured drawing. The right way to read these images is as faithful concept renderings of a form you authored - excellent for discussing material, light and mood on a committed scheme, and worthless as construction information. Keep the accuracy where it belongs, in your CAD or BIM model; let ControlNet do what it is good at, which is making that model beautiful and legible fast.
Try it

Do it yourself

Check your grasp of control, not just tools.

  1. 1What does denoise strength control in img2img, and what happens at the low and high ends?
  2. 2Which ControlNet type would you use for a grey massing model, and which for a hand sketch? Why?
  3. 3Why does ControlNet keep your composition but not your dimensions?
  4. 4You have a committed massing and want to test three materials. Where on the control dial do you sit, and why?
  5. 5Why is a ControlNet output more clearly 'your design' than a text-to-image render?
Take this with you

The one line to carry out

img2img and ControlNet constrain image AI to geometry you supply - so the AI renders your composition in new material, light and mood instead of inventing its own. You keep the form; the AI dresses it. Sit tighter on the dial the more you have committed.
Take it further
References & further reading

Peer-reviewed journals & authoritative standards

  1. 01Stable DiffusionWikipedia, 2026.
  2. 02Diffusion modelWikipedia, 2026.
  3. 03Text-to-image modelWikipedia, 2026.
  4. 04Architectural renderingWikipedia, 2026.
  5. 05Stability AIStability AI, 2026.
Related lessons
Recap
When you have a form you believe in, plain text-to-image is the wrong tool. img2img starts from your image and drifts as far as the denoise dial allows; ControlNet goes further, conditioning Stable Diffusion on a structural map - Depth for massing, Canny or Scribble for sketches - so your geometry is locked while the prompt supplies skin and light. Choose your point on the control spectrum by how much of the design you have decided, and remember these tools hold composition, not measurements: they are fast concept renderers of a form you authored, not construction drawings.
Carry forward →

With geometry under control, the exploration shifts from form to feeling. Next: using AI for moodboards, style and material studies - fast, divergent visual thinking about how a space should look and feel.

A

The author

Amogh N P

Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.

More about Amogh →