Skip to content
ROLearn Intelligence
Method

Research methodology

Where the numbers come from, how they are computed, and what they cannot tell you.

Last updated

This page describes the standing method. Each individual piece states the specific question, sample, and window it used, in the apparatus block at the top of the piece. Where a piece departs from what is written here, the piece says so and explains why.

Data sources

First-party collection

ROLearn operates a continuous collection pipeline against public platform endpoints and against telemetry that participating studios send us directly. For this publication, the relevant streams are:

  • Concurrent players (CCU), sampled on a fixed interval per game and stored as a time series. Reported figures are means over a stated window, never a single peak, unless the piece says peak.
  • Catalogue and item data for virtual goods, including price and availability over time.
  • Earned media mentions across YouTube, TikTok, Twitch, X, Instagram, Facebook, and news, discovered by keyword and by channel, then attributed to an activation.
  • Studio telemetry, where a studio has explicitly shared it. Always aggregated, never identifying an individual player, and never used to identify that studio without written permission.

Platform and third-party data

Where a platform publishes a figure, we use it and cite it. Where a third party publishes an estimate, we label it an estimate and attribute it. We do not blend an estimate into a first-party figure to produce a third number with no clear provenance.

Definitions we hold constant

Activation
A branded presence inside a game or virtual world: a branded world, an item drop, an event, a sponsored integration, or a persistent placement.
Attention
Time in the presence of the activation, not impressions. Where only impressions are available, the figure is labelled impressions.
Earned media value (EMV)
The cost of buying equivalent reach and engagement at market rates, computed per platform and per format, then adjusted for invalid traffic, incrementality, and overlap between platforms. Every EMV figure we publish states the rate card and the adjustments applied.
Retention window
Days since a player's first session in the measured game, not calendar days since launch.

How a figure gets computed

  1. Define the population. Which games, which activations, which window. Stated in the piece, including what was excluded and why.
  2. Filter synthetic traffic. Status probes, bots, and internal test accounts are removed before any customer-facing aggregate is produced.
  3. Normalise. Where two sources measure the same thing differently, both are converted to a stated common definition before comparison, or they are not compared.
  4. Aggregate. Medians are used for skewed distributions, which most game metrics are. Where a mean is used, the distribution is shown or described.
  5. State uncertainty. Sample size is always given. Where an interval can be computed, it is given. A point estimate with no sense of its spread is not a finding.

Known limitations

These apply to essentially everything published here, so they are stated once:

  • Public endpoints bound the sample. Platform APIs rate-limit and paginate. Coverage is therefore a large sample, not a census, and the piece says how large.
  • Social discovery is keyword-bounded. Mentions that use neither the brand term nor a tracked channel are invisible to the pipeline. This biases toward well-named activations.
  • Some platforms cannot be searched at all through official APIs, so coverage there depends on scraped sources with their own gaps. Pieces relying on those platforms say so.
  • Attribution is inference. Correlating a mention or a visit spike with an activation is not proof it caused it. Where we claim causation, we say what design supports it.
  • Historical data is retained for a finite window at full resolution and rolled up after that, so very long look-backs use coarser data. The piece states the resolution it used.

Revisions

When an upstream source corrects itself, or when we find an error in a computation, the affected figures are recomputed and the change is logged on the corrections page. Published numbers are never edited silently.

Reproduction and data requests

If you want the underlying series behind a chart, ask. We share what licensing and studio agreements allow, in CSV, with the same definitions used in the piece. Write to intelligence@rolearn.dev.