Studio Matrx Monthly · Volume 1 · Issue 4 · September 2026
Amogh N P
 In loving memory of Amogh N P — Architect · Designer · Visionary 
Evaluating Generated OptionsLesson 3.4
Generative & Parametric Urbanism/Module 3 · Generative Urban Design

Lesson 3.4 · Generative Urban Design

Evaluating Generated Options

Generation is easy and abundant; the decisive, underrated discipline is evaluation - judging among thousands of options through analysis, optimization and human and political judgement - and the deepest truth of the module is that the evaluation criteria carry all the values, so what you choose to score is what you get, and the criteria must be explicit, contestable and democratically set

12 min Interactive lessonFree · open lessonByAmogh N P· Architect & interior designer
The hook

Generating a thousand cities is easy. Knowing which one is good - and by whose values - is the whole game.

Here is the twist that reorganises everything you have learned in this module. All the generative power - the rules that grow street networks, the grammars that carve blocks, the AI that produces plausible fabric - makes *generating* options astonishingly easy. Press the button and a thousand candidate masterplans pour out. And that turns out to be the easy part. The hard part, the part that actually decides whether generative urbanism helps or harms, is evaluation: judging, among all those options, which few are actually good - and worth letting anywhere near a real city.

This is the discipline the whole field lives or dies by, and it is routinely underrated because generation is so seductive. Anyone can produce a thousand options; the skill is in the judging. And judging is genuinely hard for a reason that goes to the heart of the course: the criteria you evaluate by carry *all the values*. Choose to rank options by density and you will crown a different city than if you rank by walkability, or by minimising displacement - the same thousand options, three different winners, three different futures. So the evaluation criteria are not a technical afterthought; they are where the politics, the equity and the meaning of the whole exercise are decided. Get evaluation right - honest about what it can measure, humble about what it cannot, explicit about whose values it encodes - and generative urbanism can genuinely serve a city. Get it wrong and all that generative power just builds the wrong thing faster.

Generation is EASY (a thousand options). Evaluation is HARD and decides everything. Three lenses: analysis + optimization INFORM, human/political judgement DECIDES. The criteria carry ALL the values: what you score is what you get. Make them explicit + democratic.

Generation is easy, evaluation is hard

The first thing to internalise is the asymmetry: generation is easy, evaluation is hard, and the whole weight of generative urbanism rests on the hard side. Modern methods make producing candidate forms cheap, fast and practically unlimited - procedural rules, optimization search and AI can throw off hundreds or thousands of plausible masterplans, street networks or massing schemes with little effort. That abundance feels like progress, and it is genuinely useful for exploration. But it quietly relocates the entire problem. When options are scarce, having more is valuable; when options are effectively infinite, the scarce and decisive resource becomes the ability to tell which ones are actually good. The bottleneck moves from making to judging.

And judging a thousand urban options is far harder than it sounds. You cannot meaningfully inspect them all by eye, so you are pushed to reduce each to numbers you can rank - which immediately privileges the measurable over everything else and drags you toward the optimization trap. The options are genuinely complex, each strong on some dimensions and weak on others, so there is rarely a clean winner. And the qualities that matter most - whether a place will feel alive, whom it serves, whether it is just - are exactly the ones hardest to assess at a glance or at scale. So the abundance of generation, far from solving design, sets the real problem: an evaluation problem, and a hard one.

This reframes what competence in generative urbanism actually is. It is tempting to think the skill is in the generation - the clever rules, the powerful model, the impressive quantity of output. But producing options is the commodity; discernment is the craft. The urbanist who matters is not the one who can generate the most schemes, but the one who can look at a vast generated field and judge well - knowing what to measure and what cannot be measured, which trade-offs are acceptable, whose interests each option serves, and when to throw the whole batch out because the goals that generated it were wrong. Generation gives you a haystack; the entire value is in the judgement that finds, or refuses, the needle. Everything that follows in this lesson is about doing that judgement honestly.

GENERATION IS EASY - EVALUATION IS HARDGENERATEcheap, fast, abundantthousands of optionsEVALthebottlenecka few worth havingjudged, not just rankedwhich areactuallyGOOD?anyone can make a thousand options; the whole discipline is judging which few deserve to exist
Zoom
The asymmetry that reorganises the module. Modern methods make generating options cheap, fast and effectively unlimited - a thousand plausible schemes at the press of a button. That feels like progress, but it relocates the whole problem: when options are effectively infinite, the scarce and decisive resource is the ability to judge which few are actually good. Evaluation, not generation, is the bottleneck. Producing options is a commodity; discernment - knowing what to measure, what cannot be measured, and when to reject the whole batch - is the craft. A haystack is not an answer.

Press the button -> a thousand options (cheap, abundant). The real bottleneck is EVALUATION: which few are actually good? Generation is a commodity; discernment is the craft. A haystack is not an answer.

Three lenses - analysis, optimization, human and political judgement

How do you actually evaluate a field of generated options? Through three complementary lenses, none sufficient alone, in ascending order of what they can see. The first is analysis: measuring each option on quantifiable dimensions - density, daylight, connectivity, travel time, cost, walkability or space-syntax scores. Analysis is genuinely powerful and the natural partner of generation, because it can be applied at scale: you can compute these metrics across a thousand options automatically, which is the only way to handle that many. It grounds judgement in evidence rather than impression. But analysis sees only the measurable, so on its own it will quietly redefine 'good' as 'high-scoring', which is the trap.

The second lens is optimization: using the metrics not just to describe options but to *rank and search* them - steering the generator toward high-scoring regions, or picking out the options that best balance competing goals along a trade-off or Pareto front (the subject of Module 5). Optimization is how you make sense of thousands of options against multiple objectives, and it is genuinely useful for narrowing a huge field to a shortlist worth human attention. But optimization is only ever as good as the goals you gave it: it finds the best option *by your metrics*, which is not the same as the best city, and it can converge hard on the measurable while the unmeasurable is sacrificed. It informs; it does not decide.

The third lens is the decisive one: human and political judgement. This is where the qualities analysis and optimization cannot see re-enter - meaning, belonging, justice, the feel of a street, and above all the questions of whom an option serves and who it displaces. It is where the affected community's voice belongs, where the trade-offs that are really value choices get made in the open, and where the binding decision is legitimately taken. The three lenses form a hierarchy of trust: analysis and optimization *inform and narrow* - they are indispensable for handling scale and grounding the conversation in evidence - but human and democratic judgement *decides*, and must never be skipped or replaced by a score. Skip the third lens and you have automated the optimization trap; keep it, and computation becomes a genuine servant of a human choice.

THREE LENSES - NONE SUFFICIENT ALONE1. ANALYSISmeasure each option- density, daylight, access- connectivity, cost- walk / space-syntax scoressees the measurableblind to the rest2. OPTIMIZATIONrank / search by goals- best on the metrics- trade-offs, Pareto front- narrows the fieldonly as good as the goalsoptimizes the measurable3. HUMAN + POLITICALjudge what matters- meaning, justice, belonging- who is served / displaced- the community decidessees the unmeasurableholds the binding choiceanalysis + optimization inform; human and democratic judgement decides - and must never be skipped
Zoom
Three lenses for evaluating a generated field, none sufficient alone. Analysis measures each option on quantifiable dimensions (density, daylight, access, cost) and can be applied at scale - but sees only the measurable. Optimization ranks and searches by goals, narrowing thousands to a shortlist along trade-offs and Pareto fronts - but is only as good as the goals it was given. Human and political judgement decides, seeing what the others cannot: meaning, justice, belonging, and whom the option serves or displaces. Analysis and optimization inform and narrow; human and democratic judgement decides - and must never be skipped or replaced by a score.

The criteria carry all the values

Now the deepest point in the module, the one to carry out above all others: the evaluation criteria carry all the values. Evaluation feels like the objective, technical end of generative design - you are just measuring and ranking, after all. It is the opposite. The moment you choose what to measure and how to weigh it, you have made a value-laden, often deeply political decision about what a good city *is* - and that decision, not the generation, determines the outcome. What you choose to score is what you get.

Make this concrete. Take the same thousand generated options and rank them by three different criteria. Rank by maximum density and the winner is a scheme of packed towers. Rank by maximum walkability and the winner is a fine-grained, low-rise fabric. Rank by minimum displacement and the winner is the one that keeps the existing informal settlement intact. Three criteria, three completely different winning cities, three different futures for the people who live there - and the generation was identical. The criteria did all the deciding. So when someone presents 'the best option' from a generative process, the only question that matters is: best *by what criteria, chosen by whom, and for whom*? Hidden inside every evaluation is an answer to 'what kind of city, for which people', and whoever controls the criteria controls the result while appearing merely to measure.

This is why evaluation is the ethical and political core of generative urbanism, not its neutral back end. Bury the criteria in a scoring function and you launder a political choice as a technical one - 'the analysis shows this option is best' - exactly the 'the algorithm says' move the whole course warns against. The disciplined response follows directly. Make the criteria explicit: state plainly what is being optimized and what is being ignored. Make them contestable: expose them to challenge, especially by the affected communities, because reasonable people weigh density against displacement differently and that weighing is a public question. And set them democratically: the criteria for judging a city's future are themselves a matter for the planning authority, the participatory process and the community, not for the analyst to choose quietly. In generative urbanism, whoever writes the evaluation criteria writes the city - so the criteria must be dragged into the light and decided in the open.

THE CRITERIA CARRY ALL THE VALUESsame 1000 generatedoptionscriteria: max DENSITY-> winner: packed towersa different citycriteria: max WALKABILITY-> winner: fine-graineda different citycriteria: min DISPLACEMENT-> winner: keeps the informala different citywhat you choose to score is what you get - the criteria ARE the politicsso the criteria must be explicit, contestable and set by the democratic process, not buried in a score
Zoom
The deepest point of the module: the evaluation criteria carry all the values. Take one identical set of a thousand generated options and rank it three ways. By maximum density, the winner is packed towers; by walkability, a fine-grained low-rise fabric; by minimum displacement, the scheme that keeps the existing informal settlement. Three criteria, three completely different winning cities - and the generation never changed. What you choose to score is what you get, so the choice of criteria is a political act, not a neutral measurement. The criteria must be explicit, contestable and set by the democratic process, never buried inside a score.

Same 1000 options: rank by density -> packed towers; by walkability -> fine grain; by minimum displacement -> keep the informal. Three winners, three cities. The criteria did all the deciding. So make them explicit, contestable, democratic.

The module in one stance - explore by machine, decide by community

Pull the module together into a practical stance. Generative methods - procedural, search-based and AI - are genuinely powerful for one thing above all: exploring the space of the possible, throwing off many options to think with, breaking fixation and surfacing the non-obvious. But generation is the easy, abundant part, and it decides nothing. The decisive discipline is evaluation, and evaluation must be done with all three lenses in their proper order: analysis and optimization to handle scale and ground the conversation in evidence, and human and political judgement - with the affected community - to actually decide, because only it can see the unmeasurable and hold the questions of justice and belonging that make a city worth living in. And running through the whole thing is the truth that the evaluation criteria carry all the values, so they must be explicit, contestable and democratically set.

Hold, too, the honest limits stacked across the module. The generated space is only as rich as the encoding, so the best real option may sit outside it. Procedural plausibility is not goodness, and its games-and-film origin leaves it blind to the owned, lived-in, informal city. AI's plausible surface can be confidently wrong and carries the biases of its training data, erasing the informal city by omission - acute in India, where so much of the city never enters a dataset or a tidy parcel. None of these are reasons to refuse computation; they are reasons to wield it with discipline and to keep human judgement firmly in charge.

So the boundary, one last time and firmly. Generative urban design helps you explore, analyse and choose among possibilities; it does not decide the city. The actual planning and land-use decisions, the statutory approvals, and the social, equity and political judgements about a city's future - including the criteria by which options are judged and the decisive question of whom the result serves - belong to the planning authority, the democratic and participatory process, the affected communities, and the governing planning law and development-control regulations, in India the master-plan and development-plan process, the applicable DCR and the National Building Code of India. Generation is a way of thinking harder about what a place could be; evaluation, done honestly and democratically, is how a community decides what it should be.

Verify-this: evaluation decides, not generation; and the criteria carry the values

Generation easy, evaluation hard

Where the real work is

Modern methods make options cheap and abundant; the scarce, decisive skill is judging which few are good. Discernment is the craft, not output. Modules 3.4, 3.1.

Three lenses of evaluation

How to judge a field

Analysis measures at scale; optimization ranks and narrows by goals; human and political judgement decides, seeing the unmeasurable and the questions of justice. Analysis and optimization inform; judgement decides. Modules 3.4, 5.1, 6.4.

The criteria carry all the values

The deepest point

The same options ranked by density, walkability or displacement crown different cities - what you choose to score is what you get. The criteria must be explicit, contestable and democratically set, never buried in a score. Modules 3.4, 9.4, 5.4.

Explore by machine, decide by community

The module's boundary

Generative methods explore the possible; the binding choice - and the criteria behind it - belong to the planning authority, the participatory process, the communities and the law: in India the master-plan process, the DCR and NBC India. Modules 3.4, 7.3, 7.4.

Hands-on workshop

Workshop - change the criteria, change the city

The module's central truth becomes unforgettable when you feel it directly: that the evaluation criteria, not the generation, decide the result. In this workshop you will take one small set of options and watch the winner change completely as you change what you value - then confront who should choose the criteria.

Just a handful of options and a notebook. No software needed to prove that criteria decide - and the binding planning, land-use and equity decisions, including the criteria themselves, always stay with the planning authority, the affected communities, the democratic process and the law.

Given & goal
Goal: prove to yourself that the criteria carry the values
Inputs: 5-6 quick, different layout options for a small site (sketch them, or reuse the ones from lesson 3.1) + a notebook
Time: ~45 minutes
  1. 1Lay out the field: put your 5-6 options in a row. Give each a rough profile - is it dense or loose, walkable or car-scaled, does it keep or clear any existing/informal fabric, who could afford to live in it?
  2. 2Rank three ways: score and rank all the options first by maximum density, then by walkability, then by minimum displacement of existing residents. Write down the winner for each.
  3. 3Confront the result: note whether the three winners are the same or different. Almost certainly the criteria, not the options, decided - write one line on what that means.
  4. 4Ask the political question: for each criterion, write who benefits and who loses if it wins. Then ask: who, in a real project, should get to choose which criterion rules - the analyst, the developer, the authority, or the affected community?
  5. 5Reflect as reasoning: a paragraph on why 'best option' is meaningless without 'best by what criteria, chosen by whom, for whom', and why the criteria must be explicit, contestable and democratically set - flagged as reasoning, not a plan.

You’ll walk away with
A one-page demonstration: your options, three different rankings with three (likely different) winners, a who-wins-who-loses note per criterion, and a reasoning paragraph on why the criteria carry the values and must be chosen democratically. Keep it - it is the module's core lesson in one page.

The worked example

Three altitudes on the same idea

Read the band that fits you — or all three.

For the architect / urban designerUsing computation to explore, analyse and test urban form - while people and the democratic process decide

For the architect or urban designer, the lesson is that your value has moved from generating options to judging them - discernment, not output, is the craft. Generative methods make a thousand schemes cheap; the scarce skill is telling which few are actually good. Evaluate with all three lenses in order: analysis to measure at scale, optimization to narrow a huge field against multiple goals, and human judgement - with the community - to decide, because only it sees the unmeasurable and the questions of who is served and displaced. Above all, treat the evaluation criteria as the design's true value-decision: the same options ranked by density, walkability or displacement crown three different cities, so what you choose to score is what you get. Make your criteria explicit and contestable rather than buried in a score, refuse to let 'the analysis shows' launder a political choice, and defer the binding planning, land-use and equity decisions - and the setting of the criteria themselves - to the planning authority, the democratic process and the affected communities.

For the planner / urbanistWhere computational methods genuinely help planning and where the city's human and political life resists them

For the planner or urbanist, evaluation is where computational urbanism becomes explicitly political, and where your responsibility is sharpest. When a generative process yields 'the best option', the only questions that matter are: best by what criteria, chosen by whom, for whom? Hidden in every evaluation is an answer to 'what kind of city, for which people' - and the criteria, not the generation, decide the result. That makes the choice of criteria a public, contestable, democratic matter, not a technical setting for an analyst to fix quietly. Your job is to drag the criteria into the light: state what is being optimized and what is ignored, expose the weighting of density against displacement to genuine public challenge, and ensure the affected communities help set the terms of judgement. Use analysis and optimization to inform and narrow, never to decide, and keep the binding choices - and the criteria behind them - with the participatory process, the communities and the law. Whoever writes the evaluation criteria writes the city; make sure that is done in the open.

For the studentHow cities can be grown by rule - and why a city is a living system, not an optimization problem

If you carry one idea out of this whole module, make it this: generation is easy, evaluation is hard, and the evaluation criteria carry all the values. It is easy to be dazzled by generative power - rules, AI, a thousand options at the press of a button. But producing options is a commodity; the real, scarce, decisive skill is judging which few are actually good. That judgement uses three lenses: analysis (measure the options), optimization (rank and narrow them by goals), and human and political judgement (decide, seeing the unmeasurable and the questions of justice). And the deepest point: the same thousand options ranked by density, by walkability, or by minimum displacement produce three completely different winning cities - so what you choose to score IS the city you get, and the choice of criteria is a political act, not a neutral measurement. Learn to ask of any generative result, 'best by what criteria, chosen by whom, for whom?', insist those criteria be explicit and democratic, and remember the binding choice stays human. That is the mark of a genuinely literate, critical urbanist.

Misconception check

Evaluation is the objective, technical part of generative design - you generate the creative options, then just measure and score them to find the best one. The scoring is neutral maths, so it removes human bias from the final choice and tells you which option is genuinely best.

This gets the truth exactly backwards, and it is the most important correction in the module. Evaluation is not the neutral, technical back end of generative design - it is its ethical and political core, and it carries all the values. The instant you choose what to measure and how to weigh it, you have decided what a good city means, and that decision determines the outcome far more than the generation does. Make it concrete: take one identical set of a thousand generated options and rank them by maximum density, and the winner is packed towers; rank the same options by walkability, and the winner is fine-grained low-rise; rank them by minimum displacement, and the winner keeps the existing informal settlement. Three criteria, three completely different winning cities, three different futures for real people - and the generation never changed. The criteria did all the deciding. So the scoring is not neutral maths; it is a value-laden, usually political choice about what kind of city, for which people, dressed as objective measurement. Far from removing bias, hiding the criteria in a scoring function launders a political decision as a technical one - 'the analysis shows this is best' - which is precisely the 'the algorithm says' move that disguises power as objectivity. And no score can capture the qualities that matter most - meaning, belonging, justice, whom the place serves and displaces - which only human and democratic judgement can weigh. The honest practice is the opposite of the claim: use analysis and optimization to inform and narrow, but make the criteria explicit, contestable and democratically set, keep human and political judgement as the decider, and leave the binding choice - and the choice of criteria - with the planning authority, the participatory process, the communities and the law. What you choose to score is what you get, so who chooses, and how openly, is everything.
Try it

Do it yourself

No software needed - reason it through.

  1. 1Why is 'generation is easy, evaluation is hard', and how does that change what counts as skill in generative urbanism?
  2. 2Describe the three lenses of evaluation (analysis, optimization, human and political judgement) and what each can and cannot see.
  3. 3Show, with an example, how the same set of generated options produces different winners under different criteria.
  4. 4Explain the claim 'the evaluation criteria carry all the values' and why hiding criteria in a score launders a political choice.
  5. 5When someone shows you 'the best generated option', what three questions must you always ask - and where does the binding decision belong?
Take this with you

The one line to carry out

Generative methods make generating urban options easy and abundant, so the decisive, underrated discipline is evaluation - judging which few options are actually good through analysis (measure at scale), optimization (rank and narrow by goals) and, above all, human and political judgement (decide, seeing the unmeasurable and the questions of justice) - and the deepest truth of the module is that the evaluation criteria carry all the values: the same thousand options ranked by density, walkability or minimum displacement crown three different cities, so what you choose to score is what you get, the choice of criteria is a political act that must be explicit, contestable and democratically set, and the binding decision - like the criteria behind it - stays human, democratic and just.
Take it further
References & further reading

Peer-reviewed journals & authoritative standards

  1. 01Multi-objective optimizationWikipedia - Multi-objective optimization, 2026.
  2. 02Pareto efficiencyWikipedia - Pareto efficiency, 2026.
  3. 03SimulationWikipedia - Simulation, 2026.
  4. 04Participatory planningWikipedia - Participatory planning, 2026.
  5. 05Space syntaxWikipedia - Space syntax, 2026.
Related lessons
Recap
This lesson reorganises the whole module around a single asymmetry: generation is easy, evaluation is hard. Procedural rules, optimization search and AI make producing candidate urban forms cheap, fast and effectively unlimited - a thousand plausible masterplans at the press of a button - which is genuinely useful for exploration but decides nothing. The scarce, decisive resource becomes the ability to judge which few options are actually good, and that judgement is genuinely hard: you cannot inspect a thousand schemes by eye, so you are pushed toward the measurable and the optimization trap, and the qualities that matter most are the hardest to assess. So competence shifts from generation, which is a commodity, to discernment, which is the craft. Evaluation works through three lenses in ascending order of what they see: analysis measures options on quantifiable dimensions and can be applied at scale; optimization ranks and searches by goals, narrowing a huge field along trade-offs and Pareto fronts; and human and political judgement decides, because only it sees the unmeasurable - meaning, belonging, justice - and the decisive questions of whom an option serves and displaces. Analysis and optimization inform and narrow; human and democratic judgement, with the affected community, decides, and must never be skipped or replaced by a score. The deepest point is that the evaluation criteria carry all the values: the same thousand options ranked by maximum density, by walkability, or by minimum displacement crown three completely different cities, so what you choose to score is what you get, and the choice of criteria is a value-laden, political act - not neutral measurement. Hiding the criteria in a scoring function launders that political choice as technical, the 'the algorithm says' move the course warns against. So the criteria must be made explicit, contestable and democratically set, especially by the affected communities. The module's stance: explore the possible by machine, but decide the city by community - the binding choices, and the criteria behind them, stay human, democratic and just.
Carry forward →

That completes Module 3: you can now distinguish generative from parametric, follow procedural generation and AI, and above all evaluate what any of them produces. Module 4 turns to what you are actually shaping when you generate - the concrete rules of urban form: street networks, blocks and density, massing and the mixing of uses.

A

The author

Amogh N P

Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.

More about Amogh →