Lesson 6.1Lesson 6.1 · Data & Analysis
Urban Data Sources
Every map of a city is also a map of what its makers chose to count - and the fabric that houses the most people is often the fabric counted least
Every dataset of a city is a decision about what counts - and someone is always left off the map.
Data-driven urbanism sounds neutral, as if the city simply hands over the numbers. It does not. Every dataset is manufactured: someone decided what to survey, how to categorise it, what to drop as noise, and who to leave off the schedule. A satellite sees roofs but not who lives under them; a census counts households on its list but not the ones it never visited; a mobility trace records the phones that could afford data plans. The map is never the territory, and in a city the gap between them is where power lives.
This lesson is a working tour of the main sources - GIS and cadastre, OpenStreetMap, census, movement and mobility data, remote sensing, and open civic data - and an honest account of what each genuinely reveals and what it structurally cannot see. The through-line matters especially in India: a very large share of the city is informal and organic, and it slips through almost every source at once. Before you build a single model, you have to know what your evidence base is blind to - because a generative masterplan can only ever design for the city its data can see.
Six sources, each a lens that sees some things and misses others. Stack them and the blind spots line up on the informal city. Data = a made record, not the city. Ask what it cannot see; ground-truth it; the binding choice stays democratic.
The main sources and what each genuinely reveals
Start with the raw materials of data-driven urbanism, because each one has a native strength you should use and a native blindness you must remember. GIS and the cadastre are the official spatial record - plots, roads, land-use zones, ownership, infrastructure - held by the planning authority or revenue department. It is the backbone of any urban model: authoritative, structured, legally consequential. But it records the *formal* city, the one that has been surveyed and titled, and it is often out of date and patchy exactly where change is fastest.
OpenStreetMap is a crowd-sourced world map - streets, buildings, points of interest, mapped by volunteers. It is astonishingly rich, free, and in some places better than any official source; it is also uneven, because coverage follows the mappers. Well-mapped neighbourhoods tend to be wealthier, more central, more visited. A quiet informal settlement may be a near-blank on OSM not because nothing is there but because no one mapped it.
Census data counts people and households - population, age, occupation, amenities - at coarse spatial units, on a slow cycle (India's decennial census). It is the bedrock of who lives where, but it is a snapshot years old by the time you use it, it aggregates to protect privacy so fine grain is lost, and it systematically misses the mobile, the homeless and residents of settlements not on the enumeration schedule.
Movement and mobility data - mobile-phone location, GPS traces, transit-card taps, ride-hailing logs - reveal how the city actually moves, at a temporal resolution no census can match. They are among the most powerful new sources and among the most biased: they see the people who carry smartphones, hold bank-linked cards and appear in commercial datasets, and under-see everyone else. Remote sensing (satellite and aerial imagery) reveals built form, growth, green cover and heat from above, cheaply and everywhere - but it sees roofs, not lives, and cannot tell a thriving lane from a slum earmarked for clearance. Open and civic data - budgets, permits, complaints, sensor feeds - can be gold, where it exists, is published, and is clean enough to use.
The informal city, undercounted by design
Now the honest core of the lesson. Line the sources up and a pattern appears: they all tend to miss the same place, and it is the place that houses the most people. The informal city - the dense old fabric, the unauthorised colonies, the vast settlements that shelter hundreds of millions across India and the global South - falls into the gap in almost every dataset at once. It has no clean cadastre because it was never formally surveyed or titled. It is thin on OpenStreetMap because fewer volunteers map it. It is undercounted in the census because enumeration schedules and addresses do not reach it well. It is under-seen in mobility data because its residents are less likely to appear in the commercial and smartphone streams those datasets draw on. From a satellite it may read as an undifferentiated grey mass, legible as 'to be cleared' but not as a living community.
This is not a random error you can average away; it is a systematic bias that runs the same direction every time, and its direction is against the poor. A model trained or built on this evidence base inherits the blindness. A generative masterplan optimizing 'the city' is really optimizing the city its data can see - and the fabric it cannot see is precisely the fabric most at risk of being planned over, because in the model it barely exists. The danger is not merely inaccuracy; it is that invisibility in the data becomes erasure on the ground, dressed as objectivity.
The discipline this demands is simple to state and hard to practise: for every source, ask *who and what is missing, and which way does the gap run?* Treat data as a partial, made record with an author and an agenda, not as the city itself. Where a source is silent, that silence is information - usually about who did not get counted. And when the evidence base cannot see a community, the humane response is to go and look, to bring in local and participatory knowledge, and to refuse to let the map's blank spaces license decisions about the people living in them.
How bias enters, from raw signal to confident claim
Bias in urban data is rarely a single villain; it accumulates quietly along a chain, and by the end it is invisible. Follow a mobility dataset from signal to claim. The raw signal is a stream of phone pings - already biased, because it only exists for people with phones and data plans. Cleaning drops points flagged as errors, snaps traces to known roads and fills gaps with assumptions; each choice is reasonable and each bakes in a worldview about what a 'normal' trip looks like. The cleaned data becomes a metric - say, average commute time - which compresses a messy human reality into one number. That number becomes a claim: 'the city's commute is twenty minutes', quoted in a plan as fact. By the last step the phones that were never in the sample, the trips snapped onto the wrong road, and the assumptions in the fill are all gone from view. The number looks objective precisely because its manufacture is hidden.
Three biases recur and are worth naming. Coverage bias: who is in the dataset at all (smartphone owners, titled property, mapped streets) skews toward the wealthier and more formal. Categorisation bias: the messy real world is forced into tidy boxes, and whatever does not fit a box - a shop-house that is home, workshop and store at once, a street that is also a market and a living room - is mangled or dropped. Temporal bias: data ages, and a census or land-use map years old describes a city that has already moved on, especially where change is fastest.
The practical response is not to reject data - data-driven urbanism is genuinely powerful and here to stay - but to read it like a critic. Trace every confident number back up the chain to the signal it came from. Ask what was sampled, what was dropped, and what was assumed. Cross-check sources against each other and against your own feet on the ground. Report uncertainty and gaps as first-class results, not footnotes. Data can sharpen judgement wonderfully; it cannot replace it, and the moment a number stops being questioned is the moment it starts to mislead.
Building an evidence base you can trust
Put this to work as a practical method. A trustworthy urban evidence base is not the biggest pile of data; it is a triangulated one, where independent sources check each other and their gaps are mapped as carefully as their contents. Combine the official cadastre (formal structure) with OpenStreetMap (fine-grained texture), census (who lives here), mobility (how they move), remote sensing (built form and change) and civic data (how the city is run and complained about) - and treat the places where they disagree as the most interesting places, not errors to smooth over. A settlement that is dense on satellite imagery, blank on the cadastre and thin on OSM is telling you exactly where the formal record fails and the people are.
Add the source no dataset contains: ground truth and local knowledge. Walk the area. Talk to residents, local officials and community organisations. Bring participatory mapping in, so the people who live in the data's blind spots can put themselves on the map. This is not softness; it is rigour - it is how you catch the systematic errors that no amount of computation will surface, because the computer only knows what it was given. In the Indian context, where informality is the norm rather than the exception across much of the urban fabric, this ground-truthing is not optional; it is the difference between an evidence base that serves the whole city and one that quietly designs for half of it.
Finally, keep the boundary from the course's first lesson. Data informs; it does not decide. An evidence base helps you understand a place, argue from something firmer than opinion, and open options for public debate - but the binding decisions about land use, displacement and a city's future belong to the planning authority, the statutory and participatory process, the affected communities and the governing law and development-control regulations, in India the master-plan process, the applicable DCR and the National Building Code of India. The most honest thing a good dataset can do is show you clearly, and in public, both what it knows and what it cannot see.
Every source has a blind spot
What each dataset cannot see
GIS = formal only; OSM follows the mappers; census is coarse and slow; mobility sees smartphones; remote sensing sees roofs. Ask of each: who and what is missing? Modules 1.2, 6.1.
The informal city is undercounted
A systematic bias, not random noise
The sources converge on the same blind spot - the informal fabric that houses the most people - and the gap runs against the poor. Ground-truth and participatory mapping are rigour, not softness. Modules 6.1, 9.4.
Triangulate and report gaps
How to build a trustworthy base
Cross-check independent sources, treat disagreements as findings, and publish uncertainty and blind spots as first-class results. Data informs; it never decides. Modules 6.4, 1.2.
The binding choice is democratic
Who decides on land and displacement
Land-use, displacement and equity decisions belong to the planning authority, the participatory process, the communities and the law - in India the master-plan process, the DCR and NBC India - never to 'the data'. Modules 7.2, 7.3.
Workshop — audit the blind spots of one place
Pick a place you know well and interrogate what the standard data sources would and would not see of it. The aim is to feel, concretely, how the evidence base can be confident and wrong at the same time - and where its silences fall.
Free web maps (satellite plus OpenStreetMap) and a notebook. No GIS software needed - this workshop is about reading data critically by eye; the analysis tools come later, and the binding urban decisions always stay with the planning authority, the community and the democratic process.
Goal: see a real place through, and past, its data Inputs: a neighbourhood you know + free web maps (a satellite view and OpenStreetMap) + a notebook Time: ~45 minutes
- 1Choose the place: a neighbourhood, market or settlement you know on the ground, ideally one with some informal or fine-grained fabric.
- 2Look at it in OpenStreetMap and in a satellite view: note what is mapped in detail, what is a near-blank, and where the two disagree.
- 3For each source (cadastre, OSM, census, mobility, satellite), write one line on what it would show of this place and one line on what it would miss.
- 4Name the systematic gap: list the people, uses or activity that every source would tend to undercount here, and note which way the bias runs.
- 5Write a short reflection: if a generative masterplan were built only on this data, what would it get wrong about this place, and how would you bring the missing reality in - flagged as reasoning, with the binding decisions left to the process and the community.
You’ll walk away with
A one-page data audit of a real place: a per-source table of what is seen and missed, the systematic blind spot named with its direction, and a reflection on what a data-only model would erase. Keep it as a template for reading any urban evidence base.
Three altitudes on the same idea
Read the band that fits you — or all three.
For the architect or urban designer, data is the raw material of every computational move you make - and its biases become your design's biases unless you interrogate them first. Assemble a triangulated base: cadastre for the formal skeleton, OpenStreetMap for texture, census for population, mobility for movement, remote sensing for built form and change. Then, for each layer, write down what it cannot see - and notice that the answers converge on the informal city. Do not let a blank on the map read as empty ground; a generative model will happily optimize over a settlement it cannot see. Ground-truth by walking the site and bringing in participatory and local knowledge. Use data to explore and to argue, defer the binding land-use and displacement decisions to the planning authority, the participatory process, the affected communities and the governing law, and always publish what your evidence base is blind to alongside what it shows.
For the planner or urbanist, the evidence base is where objectivity is most claimed and most dangerous, because a biased dataset can dress a political choice as a technical fact. GIS, census, mobility and remote sensing genuinely strengthen how you understand a place and argue from evidence - but their gaps run in a consistent direction, against the informal, the mobile and the poor. A commute figure, a density map or a 'vacant land' classification can each quietly erase the people the data never counted. Treat every source as a made, partial record; triangulate; and report gaps and uncertainty as first-class findings that open debate rather than close it. Bring participatory mapping and community knowledge into the base so the undercounted can appear. Keep the binding decisions with the statutory process, the communities and the law - and never let 'the data shows' stand in for a choice about who the city is for.
Learn to read a dataset the way you would read a source in history: ask who made it, why, what they counted and who they left out. The main urban sources - GIS and cadastre, OpenStreetMap, census, mobility, remote sensing, open data - each reveal something real and hide something else, and their blind spots line up on the same place: the informal city that houses the most people. That is not random noise; it is a systematic bias that runs against the poor, and a model built on it inherits the blindness. The skill to build is not collecting more data but questioning it - tracing a confident number back to the messy signal it came from, cross-checking sources, and treating silence in the data as information about who was not counted. Data-driven urbanism is powerful and worth mastering; it is trustworthy only in the hands of someone who knows exactly what their evidence cannot see.
“More data means a more objective, more accurate picture of the city. If we just gather enough sources - satellite, mobile-phone traces, sensors, open data - we can finally see the whole city as it really is and plan from facts instead of opinion, removing bias from urban decisions.”
Do it yourself
No software needed — reason it through.
- 1List the main urban data sources and give, for each, one thing it reveals and one thing it structurally cannot see.
- 2Why is the informal city undercounted across almost every source at once, and why is that a systematic bias rather than random error?
- 3Trace how bias accumulates from a raw mobile-phone signal to a confident claim like 'the city's commute is twenty minutes'.
- 4What does it mean to triangulate an evidence base, and why are the disagreements between sources the most informative part?
- 5Why is a blank space on the map dangerous rather than neutral, and what is the humane response to it?
The one line to carry out
Peer-reviewed journals & authoritative standards
- 01Geographic information system — Wikipedia — Geographic information system, 2026.
- 02OpenStreetMap — Wikipedia — OpenStreetMap, 2026.
- 03Open data — Wikipedia — Open data, 2026.
- 04Big data — Wikipedia — Big data, 2026.
- 05Informal settlement — Wikipedia — Informal settlement, 2026.
Data describes a city as points, lines and counts - but a street network is more than a set of lines; its very shape channels movement and life. Next we turn form itself into a network and ask what its structure predicts, with space syntax.
The author
Amogh N P
Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.
More about Amogh →