Lesson 3.2Lesson 3.2 · AI for Concept & Ideation
Sketch-to-Image & ControlNet
Turning your sketches and massing into rendered concepts with img2img and ControlNet - keeping YOUR geometry while the AI explores only the look
What if the AI kept the building you drew, and only changed how it looks?
Plain text-to-image is a slot machine: it invents a new form every run, and none of them is the one you carefully composed. That is fine for scattering early ideas - but useless the moment you have a massing you believe in and want to see it in different materials and light.
That is what img2img and ControlNet are for. They feed your drawing into the model as a constraint, so the AI keeps your geometry - your roofline, your window rhythm, your plan - and explores only the surface: material, mood, atmosphere. This is the lesson where AI rendering stops overwriting your design and starts serving it.
ControlNet = AI as a renderer of YOUR design, not a generator of its own. Keep the form; explore the skin.
Why constraint is the whole point
In the last lesson the AI decided the form and you reacted. That is perfect for opening the solution space, but the opposite of what you want once you have committed to a composition. When you have a massing model, a plan, or a considered perspective sketch, you do not want the AI to reinvent it - you want it to render it. The problem is that a text prompt alone gives the model no idea of your geometry, so it will cheerfully produce a different, prettier building every time.
Image-to-image (img2img) solves half of this. Instead of starting from pure noise, the model starts from your image and partially denoises it toward your prompt. A single dial - variously called denoise strength or image weight - controls how far it drifts. Low strength (say 0.2-0.35) barely changes your input, just restyling surfaces; high strength (0.7+) keeps only a loose echo of your composition and invents the rest. Learning to feel that dial is most of img2img.
The honest limit: img2img respects colour and rough layout, but under a strong prompt it will still bend edges, move openings and hallucinate detail. For tight control - where a roofline or a column grid must survive exactly - you need something that reads structure, not just pixels. That is ControlNet.
It is worth being clear about why you would want constraint at all, because it feels like giving up the AI's magic. The answer is ownership. A text-to-image concept, however beautiful, is a suggestion the model made; a constrained render is your design, seen in a new light. As a project matures, the second is almost always what you need - to show a client the scheme you agreed, to test materials on the massing you resolved, to keep a competition entry recognisably one building across ten images. Constraint is not a limitation on creativity; it is what lets you apply the AI's speed and beauty to a decision you have already earned, instead of endlessly re-opening it.
img2img dial: low denoise = restyle only; high denoise = keep a loose echo and invent the rest.
ControlNet: locking the geometry you drew
ControlNet is an add-on to Stable Diffusion that conditions the image on a structural map extracted from your input, not just its colours. It comes in flavours, each reading a different kind of structure. Canny and Scribble follow edges and lines - ideal for a hand sketch or a line export from CAD. Depth reads near-versus-far - perfect for a grey massing model or a SketchUp view, so the AI understands the volumes. MLSD snaps to straight lines, which keeps architectural geometry crisp. Segmentation maps regions (wall, glass, sky, ground) so you can say what each area is.
The workflow is straightforward in ComfyUI or Automatic1111: load your sketch or massing render, pick a ControlNet type, and it produces a control map the model must honour while your text prompt supplies material, light and mood. Because the geometry is locked to your input, you can now run ten variations - warm timber at dusk, white render at noon, weathered concrete under cloud - and every one keeps your building.
Input: grey SketchUp massing, eye-level view
Control: Depth (reads the volumes)
Prompt: rammed-earth facade, deep window reveals, warm
late-afternoon light, architectural photograph,
no people, no text
Result: your massing, now rendered - same form, new skinThe key mental shift: ControlNet makes the AI a renderer of your design, not a generator of its own. The composition is yours; the AI only dresses it.
A few practical notes save frustration. The quality of the control map matters: a clean line drawing or a clear grey massing with readable depth gives the model something firm to hold, while a muddy, cluttered input produces muddy, uncertain output. You can dial each ControlNet's strength or weight - how strictly the model must obey the map - so a lower weight lets the prompt breathe while a higher one clamps the geometry hard. And you can be selective about which flavour to trust: Depth is forgiving and great for volumes but vague about edges; Canny is precise about lines but ignores depth. Choosing the map that matches what you most need to protect - the silhouette, the volumes, the window grid - is half the craft of getting a faithful render out of your own sketch.
Choosing where to sit on the control dial
There is a spectrum from loose to tight, and the right point depends on how much of the form you have already decided. Pure text-to-image is loosest - the AI owns the form; use it to diverge when you have nothing committed. img2img at low strength keeps your rough layout and colour, good for quickly reskinning a concept. img2img at high strength hews close to your input. ControlNet is tightest, pinning edges or depth exactly - use it when the geometry is sacred.
A practical rule: the more of the design you have committed to, the tighter you should go. Early, go loose and let the model surprise you. Once you have a massing you believe in, go to ControlNet-Depth so exploration cannot quietly erode the decision you made. Sitting at the wrong point wastes effort - loose control on a committed design gives you pretty images of the wrong building; tight control too early kills the surprise you were hoping for.
You can also stack constraints. Combining a ControlNet map with a modest img2img strength, or two ControlNets (Depth for volume plus Canny for edges), gives fine-grained hold - though every added constraint narrows variety, so add them only as far as the design demands. As always, more scrutiny where the stakes are higher: a mood study can drift; a client-facing image of an agreed scheme should keep the scheme.
There is a natural progression across a project's life that maps onto this dial. In week one you are at the loose end, letting text-to-image throw forms at you. By schematic design you have a massing, so you slide toward img2img and light ControlNet to reskin it. By the time you are presenting a resolved scheme you are at the tight end, ControlNet-Depth locking the building while you explore only material and light. Reading where a project sits and picking the matching control level is a judgement in itself - and getting it wrong is one of the commonest ways designers waste hours with these tools, generating gorgeous images of a building nobody asked for.
More of the design decided = go tighter. Committed massing? ControlNet-Depth so exploration cannot erode it.
Where this fits your process - and where it stops
This technique sits in the fertile middle of a project: past pure ideation, before real visualization. It shines at taking a rough SketchUp or Rhino massing and making it legible and atmospheric enough to discuss - with a client, a team, a jury - without the cost of a full render. It is also a fast way to test material and light options against a fixed form: same geometry, ten skins, in the time a single V-Ray render would take.
But keep the boundaries honest. ControlNet holds your composition, not your dimensions - it will not keep a window at exactly 1200mm or guarantee a buildable detail, and it still invents texture and small features. The output is a persuasive concept image, not a construction document and not a substitute for a measured drawing. Where accuracy is contractual - a planning submission, a heritage context - a diffusion image is evidence of intent, not fact.
And the authorship principle from the last lesson still holds, now with a stronger foundation: because the geometry started as yours, the result is far more clearly your design than a text-to-image guess ever could be. That is the quiet reason this workflow matters - it lets you use the speed and beauty of diffusion while the building on the screen remains the one you drew. Module 5 picks this up for finished visualization; here it stays a concept tool.
One workflow habit ties it all together: keep the source of truth in your CAD or BIM model and treat every diffusion image as a disposable view off that model. When the design changes, you change the model and re-export the control input, rather than trying to edit the AI image. That discipline stops the common trap of a beautiful render quietly becoming the design of record - a picture nobody can dimension, coordinate, or build from. Used this way, sketch-to-image is a fast, honest window onto a design that lives properly elsewhere, and that is exactly the role it should keep.
img2img
Generating from your image plus a prompt, not from noise
Denoise strength dials how far it drifts from your input - low restyles, high reinvents. The gentlest form of control.
ControlNet
Conditioning Stable Diffusion on a structural map from your input
The core tool for keeping your geometry. Comes in types (Canny, Depth, MLSD, Segmentation) for different inputs.
ControlNet Depth
Reads near/far from a massing model or SketchUp view
Best for turning grey volumes into rendered concepts - the AI understands the 3D form, not just outlines.
Canny / Scribble
Edge- and line-following ControlNet types
Ideal for hand sketches and CAD line exports where the outline must be honoured exactly.
Denoise strength
The img2img dial for how much the output deviates
The single most important lever in sketch-to-image; feel it out on a fixed input before trusting it.
Workshop — render a form you drew
You will take one piece of geometry you authored - a hand sketch, a CAD line drawing, or a grey massing render - and produce several rendered concepts that all keep that form, using img2img or ControlNet. The aim is to feel the control dial and prove the building on screen is still yours.
Stable Diffusion with the ControlNet extension (ComfyUI or Automatic1111 - free, local) is ideal; a hosted img2img tool works for the img2img steps. Plus your own sketch or massing view.
Goal: get multiple material/light options of ONE form you designed Inputs: a sketch or grey massing image + Stable Diffusion with ControlNet (ComfyUI/Automatic1111), or an img2img tool Time: ~45 minutes
- 1Prepare one input image: a clean line sketch, a CAD edge export, or a grey SketchUp/Rhino massing view (eye-level reads best).
- 2Start with img2img: set denoise low (~0.3) with a material+light prompt, then rerun at ~0.6. Note how much your form survives at each.
- 3Switch to ControlNet - Depth for a massing, Canny/Scribble for a sketch - and generate with the same prompt. Compare how much better it holds your geometry.
- 4Hold the ControlNet input fixed and change only the prompt three times (three materials or three times of day). Confirm every output keeps your form.
- 5Pick the strongest image, and in one sentence state what stayed yours (composition) and what the AI supplied (skin, light, mood).
You’ll walk away with
A small comparison sheet: your input, one img2img result at low and high denoise, and three ControlNet variations of the same form - with a note on which control level held your design best and why.
Three altitudes on the same idea
Read the band that fits you — or all three.
This is the bridge from your massing model to a conversation-ready image. Export a grey SketchUp or Rhino view, run ControlNet-Depth, and show a client the agreed scheme in three material and light options in an afternoon - no render farm. Use Canny or MLSD from a CAD line export when the elevation geometry must stay crisp. Just remember it holds composition, not dimensions: it is a concept renderer, not a drawing.
Photograph or model the actual room, then restyle it against locked geometry. img2img at low denoise, or ControlNet from a perspective sketch, lets you show a client THEIR space in a new palette, finish or lighting mood while the walls, windows and proportions stay put. It is far more convincing than a generic concept because it is recognisably their room - a powerful, honest way to present material and mood options.
ControlNet is how you make studio work look resolved without faking the design. Instead of a generic AI render that a tutor can tell you did not author, run your own massing or hand sketch through Depth or Scribble - the form is provably yours, the AI only renders it. Learn the denoise dial and the ControlNet types now; they are among the most employable diffusion skills, and they keep your voice in the image.
“img2img and ControlNet let AI render my design accurately, so I can trust the geometry in the output.”
Do it yourself
Check your grasp of control, not just tools.
- 1What does denoise strength control in img2img, and what happens at the low and high ends?
- 2Which ControlNet type would you use for a grey massing model, and which for a hand sketch? Why?
- 3Why does ControlNet keep your composition but not your dimensions?
- 4You have a committed massing and want to test three materials. Where on the control dial do you sit, and why?
- 5Why is a ControlNet output more clearly 'your design' than a text-to-image render?
The one line to carry out
Peer-reviewed journals & authoritative standards
- 01Stable Diffusion — Wikipedia, 2026.
- 02Diffusion model — Wikipedia, 2026.
- 03Text-to-image model — Wikipedia, 2026.
- 04Architectural rendering — Wikipedia, 2026.
- 05Stability AI — Stability AI, 2026.
With geometry under control, the exploration shifts from form to feeling. Next: using AI for moodboards, style and material studies - fast, divergent visual thinking about how a space should look and feel.
The author
Amogh N P
Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.
More about Amogh →