Home

Data eng

Steam Explorer: Geospatial Game Search on MongoDB

GitHub ↗

Steam Explorer landing page showing a grid of game cards with cover image, publisher, rating, price, release date, average playtime and review count
The app: 3,230 games with ratings, playtime and review counts, all served from MongoDB.

Team project for Non-Relational Data Management, a graduate course at RIT.

games
3,230
reviews
45,501
geospatial query types
2

The idea

Game data is naturally messy and nested: every title has its own set of tags, a publisher with a real-world address, thousands of reviews and a cover image. This project collects that data, models it in a document database and puts a search app on top, including a question a normal store page cannot answer: which games were made near here?

Getting the data

  1. 01

    Collect. Pulled Steam's top-rated games (those with 5,000+ reviews) and their reviews using SteamDB's APIs and Selenium automation scripts.

  2. 02

    Clean. Merged records from several Steam sources so each game appears exactly once.

  3. 03

    Enrich. Looked up each publisher's address and converted it to map coordinates, stored as a GeoJSON point inside the game record. Without this step no location search is possible.

  4. 04

    Load. Imported the result into MongoDB as a Games collection and a Reviews collection, with cover images loaded into GridFS.

How the data is modelled

Three stores, linked by simple keys. Everything a game card needs sits in one document, so the landing page reads a single collection.

MongoDB

MongoDB

Games

One document per game

3,230 documents

  • appid · unique game id
  • publisher.location · GeoJSON point, 2dsphere index
  • tags · nested tag counts
  • icon · cover image filename

Reviews

One document per review

45,501 documents

  • appid · links to its game
  • playtime_forever · minutes played
  • location · GeoJSON point
  • timestamp_created · for sorting

GridFS

Cover images, inside the database

3,187 images

  • fs.files · filename and type
  • fs.chunks · binary pieces
Reviews point to Games through appid; a game's icon field names its image in GridFS.

A game document, trimmed:

{
  "appid": 620,
  "name": "Portal 2",
  "rating": "97.72%",
  "genre": "Action, Adventure",
  "publisher": {
    "name": "Valve",
    "location": {
      "City": "Bellevue",
      "Country": "United States",
      "location": {
        "type": "Point",
        "coordinates": [-122.2013196, 47.6140803]
      }
    }
  },
  "tags": { "Platformer": 7477, "Puzzle": 7394, "Co-op": 6260 },
  "icon": "img_620.jpg"
}

The queries behind it

  • Which games were made nearest to this point?

    $geoNear

    An aggregation pipeline that starts with a geo stage on the publisher's location. It returns games ordered by distance, includes the distance itself, and applies the name or genre filter inside the same stage so paging stays correct.

  • Which games were made within this radius?

    $geoWithin

    A circle drawn on the sphere around the chosen point. Unlike the nearest-first query, results can still be sorted by rating, name or release date.

  • How long do people really play each game?

    $match → $group

    For the games on the current page, a pipeline over the Reviews collection averages recorded playtime per game, so each card shows a figure computed from reviews, not a stored number.

Location search panel with a map, a pin and a radius circle over western Europe, and options for any location, within radius or nearest first
Location search: click a point on the map, set a radius, and choose within-radius or nearest-first.

The hard part

Location search depends on data Steam does not provide. Many games had no publisher address at all, so the coordinates had to be filled in through extra API lookups before a single geo query could run. On the query side, nearest-first search has a rule of its own: the geo stage must come first in the pipeline and it fixes the sort order, so filtering and paging had to be built around it and the total count computed by a second pipeline.

What I would add next

A text index for search, in place of the current pattern match on name and genre, and a page of analytics built on aggregation pipelines: how playtime relates to rating, and which regions produce the best-reviewed games.