Lesson 10.2Lesson 10.2 · Practice & the Future
Getting Started with City Data
You do not need a city's budget, a platform licence or a data team to begin - with an open portal, a free tool and one honest question you can build a small piece of the twin this week, and learn more by doing it than any slide deck could teach
What if the barrier to learning how a city twin works was not a licence or a budget, but simply not knowing that the data is already sitting in an open portal, free to download right now?
The cities that run full digital twins spend fortunes on platforms, sensors and specialists, and it is easy to conclude that the whole field is locked behind a paywall you will never pass as a student or a small practice. That conclusion is wrong, and believing it is the single biggest reason designers stay spectators. The machinery of a full twin is expensive; the craft of working with city data is not. Much of what a twin is built from - building footprints, road networks, terrain, open imagery, administrative boundaries - is published as open data by cities and national agencies, downloadable today, for nothing.
This lesson is deliberately practical. It will not teach you to stand up a city-scale twin; it will show you how to get your own hands on real city data, open it in free tools, build a small context model of a few streets, and run one honest experiment on it - this week, with no budget. That small loop, done once with your own data, teaches more about what a twin really is (and where its data quietly lies) than any amount of reading. The goal is not a polished deliverable. It is to cross the line from reading about twins to touching the raw material they are made of, and to start building the judgement that only comes from handling real, imperfect data yourself.
The data is already in the portal, free. A portal, a free tool, one question - start this week, stay small.
Find the data - open portals are the public on-ramp
Start where the data lives: open data. A great deal of the raw material of a city twin is published by public bodies precisely so that citizens, researchers and businesses can use it. Many cities and national governments run open data portals - searchable catalogues of datasets released under open licences - and alongside them sit geospatial portals and spatial data infrastructures that serve mapped layers: building footprints, road and rail networks, land use, administrative boundaries, terrain models, sometimes live feeds like transit positions or air quality. There are also large community and global datasets (collaborative street maps, open building footprint collections) that fill gaps where official data is thin.
The practical move is to go looking for your own city's portal and see what it actually offers. Search for the city or national open-data portal; browse the geospatial or mapping section; note what exists and in what form. You will quickly learn the vocabulary of the field by encountering it: vector layers of footprints and roads, raster layers of imagery and terrain, formats like GeoJSON and common GIS files, coordinate reference systems, and metadata describing where each dataset came from and when. Even the act of browsing teaches you something twin-builders know painfully well - that data coverage is uneven, that some layers are beautifully maintained and others years stale, and that the map is never the territory.
A word on authority and honesty from the very first step. Open portal data is superb for learning, prototyping and asking design questions, but it is not always the legally authoritative record. Cadastral boundaries, survey-grade positions and official infrastructure data come from the official custodians and licensed surveyors - in India, the Survey of India and the relevant agencies - and binding uses must go to them, not to a convenient download. Get into the habit now, on your very first dataset, of recording its source, its date and its licence. That discipline - knowing exactly what you are standing on - is not bureaucracy; it is the foundation of trustworthy work, and it separates a designer who uses data well from one who is misled by it.
Free tools and a simple context model - your first few streets in 3D
You do not need an expensive platform to work with this data. A strong ecosystem of free and open-source tools exists precisely for this. Free desktop GIS software will open, inspect and analyse the vector and raster layers you downloaded - letting you see the attributes behind the pretty picture, measure, filter and combine layers. Free and low-cost 3D modelling tools, and web-based map viewers, let you visualise the result. Between them you can do a surprising amount without paying anything, and the skills transfer directly to the professional platforms.
The first real exercise is to build a simple context model: a rough 3D block model of a few streets around a site you care about. The classic move is to take a building-footprints layer (2D outlines) and a heights attribute - or a sensible assumption where heights are missing - and extrude the footprints into blocks. Drape it over a terrain layer and you have a crude but real three-dimensional model of a real place, built from real data. It will be coarse - this is a low level of detail, boxes not buildings - and that is the point. You are not making a beauty render; you are making a usable context model, and learning that the level of detail should match the question you want to ask.
Building even this small model teaches the lessons no lecture can. You will hit missing heights and have to decide what to assume - and notice how that assumption quietly shapes every later result. You will find footprints that do not match the imagery, because one layer is newer than the other. You will wrestle with coordinate systems until two layers finally line up. Each of these frustrations is a real twin problem in miniature: data is incomplete, inconsistent and dated, and a model is only ever as good as the decisions you made to patch its gaps. Keep the model small - a few streets, not a city. A small model you understand completely is worth far more, as a learning tool, than a large one you have merely downloaded.
Run one small, honest experiment
With a small context model in hand, run a single experiment - one honest question, answered as well as your data allows. The aim is not a comprehensive study; it is to close the loop from data to a question to an insight, and to feel where the answer is trustworthy and where it is guesswork.
Good starter questions are modest and spatial. A sun and shadow study: at what time does this courtyard fall into shade in winter, given the real surrounding building heights? A simple accessibility or walk-time map: what can be reached on foot in ten minutes from this corner, along the real street network? A basic visibility or overlooking check: what does a proposed upper window actually see? A rough density or land-use read: how does built footprint vary across these blocks? Each of these can be attempted with free tools on your small model, and each mirrors, in miniature, the analyses a real twin performs (Module 4). You are not producing a binding result; you are learning how the question is posed, what data it needs, and how fragile or robust the answer is.
The discipline that makes this valuable is honesty about uncertainty - the habit that will matter for the rest of your career. Every experiment rests on assumptions: the heights you guessed, the date of the footprints, the coarseness of the terrain, the simplifications in the tool. Write them down next to the result. A sun study on assumed heights is a useful hypothesis, not a fact; a walk-time map on an incomplete path network misses exactly the informal shortcuts people actually use. Stating those caveats is not weakness - it is the single habit that separates credible analysis from a misleading picture, and it is what lets you hand a result to someone else without misleading them. For anything that will inform a real decision, the experiment tells you where to look harder and which specialist or authority to consult; it does not replace them. A binding sun-on-ground or environmental result belongs to a qualified specialist with proper methods, not to your weekend model - and knowing that boundary is itself part of the skill.
Footprints + a height guess + one question = your first experiment. Write the guess next to the answer - always.
Small experiments compound into real literacy
None of this makes you a twin engineer, and it is not meant to. What it does - and this matters more - is make you fluent in the material twins are built from, so that when you sit across from the specialists you understand what you are looking at, and when you read a twin's output you know what questions to ask of it. That fluency is exactly the literate, critical presence the last lesson argued for, and the fastest way to build it is a series of small experiments, each a little braver than the last.
There is a natural progression once you have done one loop. Swap in a new data layer and see how the picture changes. Try a different question on the same model. Add a second source and discover how two datasets disagree. Share your small study with a peer and have them poke holes in it - the holes are the learning. Over a few months of modest projects, you will accumulate something rare: a working, hands-on sense of what city data is really like, where it is rich and where it is blank, and how much confidence a given result deserves. That is a genuine, demonstrable skill and, for a student, a distinctive portfolio thread almost no other design graduate will have.
Two cautions to carry forward. First, resist scope creep and the lure of the impressive. A small model you understand completely teaches more than a sprawling one you do not; the people who learn fastest stay small longest. Second, stay honest about what your experiments are - learning exercises and design hypotheses, not authoritative analyses or approvals. The official data of record, binding engineering and environmental results, and lawful handling of any data about people all remain with the custodians, qualified engineers and the governing law. Work within that boundary and your experiments are pure upside: cheap, fast, and the most effective teacher of twin-literacy there is. The Indian context sharpens both the opportunity and the cautions in specific ways - which is exactly where we go next.
Open data & spatial data infrastructure
Publicly published layers for learning and prototyping
Portals publish footprints, roads, terrain and more under open licences - ideal to learn on, but not always the legally authoritative record. Record source, date and licence. Module 3.
Free / open-source GIS & 3D tools
Opening, analysing and visualising city data
Free desktop GIS, 3D modellers and web map viewers cover the core loop; skills transfer to professional platforms. No budget required to begin. Module 5.
Official / survey / cadastral custodians
The authoritative data of record
Binding boundary, survey and infrastructure data come from the official custodians and licensed surveyors (incl. Survey of India), never a convenient open download. Module 3.
Stated assumptions & uncertainty
Honest reporting of a small experiment
Every result rests on assumed heights, data dates and tool simplifications. Writing them down is what makes an experiment credible rather than misleading. Module 9.
Workshop - your first city-data loop, start to finish
This workshop walks you through one complete loop - find data, build a small model, run one experiment, state the caveats - for a place you know. The deliverable is deliberately humble; the learning is in having touched real data end to end.
Free desktop GIS, a free or low-cost 3D tool, internet access to an open portal, and a notebook for sources and assumptions. No paid platform, no city budget.
Goal: complete one honest data-to-insight loop on a real place Inputs: a computer, internet, free GIS and a free 3D tool, and a few streets you know well Time: a few focused hours (can be split across sessions)
- 1Find a portal: locate your city or national open-data or geospatial portal. Download building footprints and a road layer for a small area you know. Record each dataset's source, date and licence before you do anything else.
- 2Open and inspect: load the layers in free GIS. Look at the attributes behind the map - do footprints have heights? Is the road network complete? Note one thing that is missing or looks out of date.
- 3Build a small context model: extrude the footprints to heights (assume a sensible height where it is missing, and write that assumption down). Keep it to a few streets. You now have a crude 3D model of a real place from real data.
- 4Run one experiment: pick a single honest question - a winter shadow study, a ten-minute walk-time map, or an overlooking check - and attempt it on your model with free tools.
- 5State the caveats: list every assumption the result rests on (guessed heights, data date, missing paths, tool simplification). Mark which parts of the answer you trust and which you do not.
- 6Write it up: one page - the place, the data and its dates, what you built, the one question, the answer, and the caveats - framed explicitly as a learning exercise and design hypothesis, not an authoritative study.
You’ll walk away with
A one-page record of one complete city-data loop: sourced and dated data, a small context model, one experiment, and an honest list of assumptions and limits. It is the first entry in a portfolio of small, credible studies.
Three altitudes on the same idea
Read the band that fits you — or all three.
Build a real context model of your next site from open data - it is the most useful and most honest on-ramp you have. Find your city or national geospatial portal, pull building footprints, roads and terrain, and extrude a rough block model of a few streets in free GIS and 3D tools. Then run one experiment that bears on your design - a winter shadow study, a visibility check, a walk-time read - and write your assumptions next to the answer. You will learn where the open data is rich and where it lies, which is precisely the judgement you need when a city twin hands you context later. Keep it small and keep it honest: this is a design hypothesis, not a binding study. Authoritative survey and cadastral data stay with the official custodians, and binding environmental and structural results stay with qualified specialists and engineers.
The same on-ramp works at the scale you care about, with the building as your context and the room as your question. You may not extrude a district, but you can pull open data about a building's surroundings - orientation, neighbouring heights, street and daylight context - and build a small model that tells you how the outside world reaches an interior: where winter sun lands, what a window overlooks, how noise or footfall sits against a frontage. It is the same loop (open data, free tool, one honest question, written-down caveats), pointed inward. And it builds the literacy to understand how a building's own data model might one day nest into a city twin. Keep any data about actual occupants out of these learning exercises, and handle real occupancy or comfort data only lawfully and with the engineers - privacy duties are sharpest where people live and work.
This is your cheapest, fastest way to turn a topic into a skill - do the loop this week. You do not need a licence, a budget or a team. Find an open portal, download real footprints and roads for a place you know, open them in free GIS, extrude a small block model, and run one experiment - a sun study or a ten-minute walk map. The frustrations you hit (missing heights, mismatched layers, stubborn coordinate systems) are not you failing; they are the real problems of city data, and wrestling them is the lesson. Write down every assumption next to every result. Do a handful of these small, honest studies and you will have something almost no design graduate has: a hands-on, critical feel for what city data is really like - a standout portfolio thread and the foundation of genuine twin-literacy. Keep it framed as learning, never as authoritative analysis.
“Getting started with city data needs an expensive digital-twin platform, proprietary software licences and access to a city's private data systems. As a student or a small practice with no budget, there is nothing meaningful I can actually do myself - I just have to wait until I work somewhere that already has a twin.”
Do it yourself
Best done at a computer, but you can reason through the choices first.
- 1Name three layers of city data commonly published on open portals that you could use to build a simple context model.
- 2What is a 'context model', and why is extruding building footprints to heights a sensible first exercise?
- 3Why should you record the source, date and licence of every dataset from your very first download?
- 4Give one small, honest experiment you could run on a few-street model, and one assumption its result would depend on.
- 5Why are open-portal data and your weekend experiment useful for learning and design, but not a substitute for the authoritative record or a qualified specialist?
The one line to carry out
Peer-reviewed journals & authoritative standards
- 01Open data — Wikipedia - Open data, 2026.
- 02Geographic information system — Wikipedia - Geographic information system, 2026.
- 03Spatial data infrastructure — Wikipedia - Spatial data infrastructure, 2026.
- 043D city model — Wikipedia - 3D city model, 2026.
Open data and free tools are universal, but the realities of city data are deeply local - and nowhere are the opportunity and the cautions sharper than in India. Next we look squarely at urban digital twins in the Indian context: the vast promise, and the informal city that the data too often renders invisible.
The author
Amogh N P
Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.
More about Amogh →