Data eng
Steam Explorer: Geospatial Game Search on MongoDB

Team project for Non-Relational Data Management, a graduate course at RIT.
- games
- 3,230
- reviews
- 45,501
- geospatial query types
- 2
The idea
Game data is naturally messy and nested: every title has its own set of tags, a publisher with a real-world address, thousands of reviews and a cover image. This project collects that data, models it in a document database and puts a search app on top, including a question a normal store page cannot answer: which games were made near here?
Getting the data
- 01
Collect. Pulled Steam's top-rated games (those with 5,000+ reviews) and their reviews using SteamDB's APIs and Selenium automation scripts.
- 02
Clean. Merged records from several Steam sources so each game appears exactly once.
- 03
Enrich. Looked up each publisher's address and converted it to map coordinates, stored as a GeoJSON point inside the game record. Without this step no location search is possible.
- 04
Load. Imported the result into MongoDB as a Games collection and a Reviews collection, with cover images loaded into GridFS.
How the data is modelled
Three stores, linked by simple keys. Everything a game card needs sits in one document, so the landing page reads a single collection.
MongoDB
Games
One document per game
3,230 documents
- appid · unique game id
- publisher.location · GeoJSON point, 2dsphere index
- tags · nested tag counts
- icon · cover image filename
Reviews
One document per review
45,501 documents
- appid · links to its game
- playtime_forever · minutes played
- location · GeoJSON point
- timestamp_created · for sorting
GridFS
Cover images, inside the database
3,187 images
- fs.files · filename and type
- fs.chunks · binary pieces
A game document, trimmed:
{
"appid": 620,
"name": "Portal 2",
"rating": "97.72%",
"genre": "Action, Adventure",
"publisher": {
"name": "Valve",
"location": {
"City": "Bellevue",
"Country": "United States",
"location": {
"type": "Point",
"coordinates": [-122.2013196, 47.6140803]
}
}
},
"tags": { "Platformer": 7477, "Puzzle": 7394, "Co-op": 6260 },
"icon": "img_620.jpg"
}The queries behind it
Which games were made nearest to this point?
$geoNearAn aggregation pipeline that starts with a geo stage on the publisher's location. It returns games ordered by distance, includes the distance itself, and applies the name or genre filter inside the same stage so paging stays correct.
Which games were made within this radius?
$geoWithinA circle drawn on the sphere around the chosen point. Unlike the nearest-first query, results can still be sorted by rating, name or release date.
How long do people really play each game?
$match → $groupFor the games on the current page, a pipeline over the Reviews collection averages recorded playtime per game, so each card shows a figure computed from reviews, not a stored number.

The hard part
Location search depends on data Steam does not provide. Many games had no publisher address at all, so the coordinates had to be filled in through extra API lookups before a single geo query could run. On the query side, nearest-first search has a rule of its own: the geo stage must come first in the pipeline and it fixes the sort order, so filtering and paging had to be built around it and the total count computed by a second pipeline.
What I would add next
A text index for search, in place of the current pattern match on name and genre, and a page of analytics built on aggregation pipelines: how playtime relates to rating, and which regions produce the best-reviewed games.
- MongoDB
- Aggregation pipelines
- Geospatial indexes
- GridFS
- Selenium
- Node.js
- Express