Get the data

Everything behind this site, in files you can open or in a live connection your assistant can query. Take the lot in one zip, take the spreadsheet, take the single table you actually need, or take none of them and point Claude at the endpoint. The licence is CC BY 4.0, which is the lawyerly way of saying go on then.

Take it all

Everything, zipped

Every table on this page, as CSV and as JSON, plus a README that repeats all of this offline and the licence. If you are not sure which of the files below you want, you want this one.

One spreadsheet

The same tables as a single workbook, a sheet each, top row frozen and filters already on. Opens in Excel, Numbers, LibreOffice and Google Sheets. For anyone who would rather sort than script.

Connect it to your AI

No file at all. The whole directory is served live over MCP, free, with no key and no account: add one URL to Claude and it can search the sources, pull a market profile and read the trends while you talk to it. Six tools, and every answer arrives carrying its date, its link and the things it is not safe to conclude.

The tables

Sources

The directory itself. Every outlet, institution, dataset, podcast, book, person and community, with what it lets you find out, who owns it, what it costs, which markets it covers, and what to watch for. This is the table most people come for.

Markets

The twelve market profiles as one file: every sourced claim under the section it belongs to, with the link it came from, and the player population figure for each market carrying the definition that says who was counted. Nested rather than flattened, because a population figure without its definition is a number waiting to be misused.

Trends

What is moving the industry, each with what it is, why it matters, the markets it lands in and how far along it is. The further reading is its own file, one row per link, so you can sort a reading list instead of unpicking a cell.

The other tables

Industry

Storefronts, consoles, livestreaming and UGC platforms, with who owns each one. Genres with their definitions, their subgenres, and the terms actually used in Chinese, Korean and Japanese. The example games are their own file, one row per game.

Searching tips

Ways of finding things that I wish someone had told me in month one. Each one says what to type, gives a query you can paste, and explains what it surfaces and why.

Vocabulary

Every value each field can hold and how many rows carry it. Read this first if you are building something on top of the data: it is the fastest way to see that Japan has 139 rows and Italy has 24 and to decide whether that gap matters for what you are doing.

How the files are shaped

One row is one thing: a source, a claim, a trend, a game. Where a row carries a list of other things, that list gets a file of its own rather than a cell with five links glued into it. That is why the further reading on a trend and the example games in a genre are separate tables, each keyed by the name of the thing they belong to.

A list of plain words stays in its cell. Markets and languages are written into one cell with a semicolon and a space between them, because splitting three tokens into a table of their own helps nobody.

CSV is UTF-8 and RFC 4180, quoted where a value needs it. Plenty of these rows carry Japanese, Korean, Chinese and Arabic, so if your spreadsheet opens one and shows you a fence of question marks, it has guessed the encoding. The workbook above is the quicker fix.

JSON is the same data with the nesting left in. There is nothing in the JSON that is not in the CSVs; it just has not been flattened, which is what you want if you are feeding this to code or to a model.

An empty cell means the field was not collected for that kind of row, not that the answer is nothing. Cost is asked of data platforms, where the price decides whether you can use the thing at all. ISBN is asked of books. Neither is missing from a subreddit.

There is no ID column. Rows are identified by name, and the row order is export order, which is not a ranking of anything by anyone. If you need keys that survive the next pass, build them yourself from the name and the link.

Not checked

These files say what a source is. They do not say whether it is right.

Every source here was opened and read before its description was written. What the source then claims is not checked. A figure you reach through this data carries exactly the authority of whoever published it, and moving it into a spreadsheet does not upgrade that. Where an owner has a stake in what gets reported, the row says so, and that is a flag to read carefully rather than a correction to the number.

Derived from The verification standard set out on Method, restated here because a file travels away from the page that explains it.

Fields in the source table

name
What the source is called. The closest thing to a key this table has. Two things can share a name, in which case the link settles it.
group
One of the six doors of the site: Data, Institutions, Press, Media, People, Communities. Every row has exactly one.
medium
What kind of thing it is, from one list of thirteen values that reads across the whole table: news outlet, organisation, forum or community, video channel, data tool, podcast, person, book or paper, company disclosure, report or document, newsletter, research firm, link list.
format
The finer classification, read inside its group. A Press row is a national daily or a trade outlet; a Data row is an analytics platform or a corporate filing. Every value, and how many rows carry it, is in the vocabulary file.
what_it_lets_you_find_out
What you can actually get from it, written after reading it, including the limits worth knowing before you cite it. The longest field in the table and the reason the table exists.
markets
Which markets it covers, as codes, listed alphabetically and meaning nothing by the order. Global means it covers the lot. Not market specific means market is not the axis the source works on.
language
The language or languages you will be reading it in. Not a claim about translation being available. Several are listed alphabetically, and that order is arbitrary: it does not say which language the site leads in.
cost
Free, Paid, Freemium or Unknown. Collected wherever the sourcing asked the question, which is every part of the table except the press outlets and the individual people, neither of which is priced. Empty means it was never asked, not that the thing is free.
link
Where to go. One address, and the front door rather than a deep link, because deep links rot first.
extra_links
Second and third addresses for rows that need them. Empty on every row today. It stays in the file because it is part of the shape rather than a promise.
owner
Who owns or runs it, where that is knowable and matters. Filled on roughly three rows in five.
ownership_note
Where the ownership changes how you should read the source: a platform that absorbed its rival, an outlet owned by the group it reports on, a body funded by the companies it represents.
beat
On people, what they cover.
where_they_post
On people, the platform you will actually read them on, which is frequently not where the link goes.
year
Publication year, on books and papers.
isbn
Thirteen digits, no hyphens, on books that have one.
audience
Who a media source is for. On video, podcast and book rows.

Fields everywhere else

Markets
One record per market. sections holds the claims, grouped under Market size, Shape, Press structure, Venues, Where the data lives, Unique to this market and Standing. Each claim carries its source and its link, plus a kind of Narrative or Card entry that says where it renders on the site rather than anything about the claim itself. player_population holds the figure as the source published it, a plain number beside it, the measure and definition stating who exactly was counted, the year, and where it came from. Ten of the twelve markets have one; MENA and India have no comparable figure and so have none. Read the definitions before you chart anything: these figures are not comparable across markets, and treating them as if they were is how bad slides get made.
Trends
slug is the trend's permanent key, and the reading file carries it as trend_slug. stage is Emerging, Dominant, Perennial or Receding. category sorts them into six groups. horizon places the trend on the Three Horizons as the trends page draws them: h1, h2, h3 or perennial. horizon_group sorts the h2 trends four ways, and horizon_line is the line written for that reading. The placements are an editorial reading, not a score. what_it_is carries the figures, why_it_matters is the short form, where_to_follow names who is covering it, and as_of is when it was last looked at. The reading file types every link as Reporting, Analysis, Data, Primary, Community or Research. The JSON also carries each trend's claims, every one with the link it came from.
Platforms and genres
Platforms carry a category, an owner and the markets they matter in. Genres carry a segment saying where the genre lives, a definition, its subgenres, and local_terms: the words used in Chinese, Korean and Japanese, which is the field that saves you when a vendor's taxonomy and yours disagree. The example games file names each game and says why it is there.
Searching tips
tip is the technique, how is the thing to type, example is a query you can paste as it stands, why explains what it surfaces, and asked_as is the question it answers in a researcher's own words. markets says where it applies.
Vocabulary
facet is the field, value is one of the values it takes, rows is how many rows carry it. within_group is filled only for format, which is read inside a group and nowhere else.

What to tell your LLM

Two ways in, and they are good at different things. Connect the directory over MCP and ask, and the model reads the rows it needs at the moment it needs them, with no file to attach and nothing to go stale on your desk. Attach the file, and you have a fixed copy you can count, check, cite and rerun next month against the same rows. Reach for the connection when you are asking; reach for the file when you are working.

Whichever you use, do not paste the data into the chat. A model handed sources.csv as an attachment can filter it, count it and quote it. A model handed several hundred rows of it as text starts summarising them, and summarising is where invented rows come from.

The prompts below all do the same two things: they name the file and the columns, and they say what the model must not do. That second half is the half that matters, and it matters over the connection too. Left alone a model will rank these sources, fill an empty cell with something plausible, or quietly refresh a 2026 figure from memory, and nothing in the data gives it a way to notice. Take one, change the market or the question, and go. If you are connected rather than attached, drop the sentence about the file and keep everything after it.

  • Start here, every time

    I have attached files from Delightful's Game Research Starter Pack, a directory of where to research the video games industry. The sources file holds 978 rows, one per source, under the columns name, group, medium, format, what_it_lets_you_find_out, markets, language, cost and link. It is called sources.csv, and if it reached you under some other name, those columns are how you know you have the right file. Treat it as a map of where to look, not as a store of facts. Every source was read before it was described, but nothing the sources themselves claim has been checked, and nothing in the file is ranked or scored, so row order means nothing. When you answer, name the rows you used and give me the link from the row. If the file does not cover what I asked, say so instead of filling the gap from what you already know. Before I ask you anything, tell me what you have: how many rows the file holds, what values the group column takes and how many rows sit under each, and how many distinct market codes appear in the markets column, bearing in mind that one cell may hold several separated by semicolons. If you cannot work those out from the file itself, say so plainly, because that tells me it did not reach you in a form you can count and everything after it would be guesswork.

  • Build a reading list for one market

    From sources.csv, give me everything covering Japan: the rows whose markets column contains JP. Group them by the group column, and inside each group put the rows whose language is Japanese first. Return a table of name, what it lets you find out, and link. Then tell me which languages I will need, and flag any group holding fewer than five rows, because that is thin coverage in this directory rather than a fact about Japan.

  • Find who publishes a number

    I need to state how much was spent on mobile games in Germany last year, and I need something citable. Using sources.csv, list every row that could plausibly publish that figure, starting with the Data and Institutions groups. For each one, tell me what the ownership_note column says where it is filled. Cost was only collected on the Data group, so it is empty on every Institutions row: where it is empty, say the file does not record it rather than reading that as free. Then order them by how likely I am to get the number without a subscription, and say where you are ordering on an empty cell. Do not give me the figure itself. I am going to go and read it.

  • Shortlist the free tools

    From sources.csv, take only the rows where group is Data and cost is Free or Freemium. Those are the two values that mean I can get something without paying, and if any row in that group carries no cost at all, list those separately as unknown rather than dropping them or ruling them in. For each, tell me what it measures, what it explicitly excludes, and which of them draw on the same underlying data so I do not check one number twice and think I have confirmed it. The what_it_lets_you_find_out column states the known limits. Use those rather than your own impression of these tools.

  • Write a trend brief

    Using trends.csv and trend-reading.csv, joined on the trend name, write a one-page brief on the trends whose stage is Dominant or Emerging and whose markets include Korea or Japan. For each one, quote the figures in what_it_is exactly as written, add the why_it_matters line, and pick the two most useful links from the reading file, naming their type. Do not add a trend that is not in the file, and do not update any figure from your own knowledge. The as_of column tells you how fresh each row is; say so where it is old.

  • Search in the local language

    From genres.csv, pull the local_terms column for shooters, role-playing games and puzzle games, and give me the Chinese, Korean and Japanese terms with a romanisation. Then, using the how and example columns of searching-tips.csv, turn each term into two or three queries I can paste straight into a search engine to find market coverage rather than store pages. Keep the local script in the queries.

  • Check the coverage before you claim anything

    Using vocabulary.csv, tell me how many rows this directory holds per market and per language. Then name the markets where the count is low enough that an absence in this data says more about the directory than about the market itself. I am about to present findings from these files and I want to know where I would be arguing from a gap.

Licence

Creative Commons Attribution 4.0. Use it, cut it up, publish from it, build something you charge for. The one condition is credit: name this site, link the licence, and say if you changed something.

The licence covers the data in these files. It does not cover what the rows point at. The reporting, figures, artwork and writing at the far end of every link belong to whoever made them, and nothing here hands you any part of that.

Working with the data

Method

Where the files came from and what was done to them. How the research ran, what Claude was and was not allowed to decide, what was verified, what deliberately was not, and the two problems in this data that no amount of care fixes.

Corrections

Something wrong, something missing, something that should not be here at all? Tell me. Corrections go into the workbook and come out in the next pass, so the files change with the pages.

What has changed

Every accepted export and every change to the site leaves a line here, newest first. The dates are the workbook's and git's, the counts are the ones the export recorded at the time, and none of it is typed by hand, so this list cannot quietly disagree with the files above it.

A directory line says what arrived, what was removed by name, and how many rows already here were corrected. A site line says what the site itself does differently. Show one kind or read both. For the story behind a correction, read Corrections.

Show
6 October 2026Directory
12 sources corrected.
1 October 2026Directory
1 source added to Data. The directory stands at 978.
30 September 2026Directory
41 sources added to Institutions (18), Press (8), Media (7), People (4), Communities (2) and Data (2). The directory stands at 977. 5 sources corrected.
28 September 2026Directory
5 sources added to Institutions (4) and Data (1). The directory stands at 936. 17 sources corrected.
28 September 2026Site
The homepage says when the directory was last updated, under the search box. Every date the site gives for its data now moves when a source or a trend does: on Method, Query it live, this page, the trends page, the downloads, the open data and the picture a shared link shows. They used to move only when the research workbook was exported, and went on saying 17 September through ten days of new sources and trend edits.
27 September 2026Directory
6 sources added to Data (3), Press (2) and Institutions (1). The directory stands at 931. 2 sources corrected.
26 September 2026Directory
1 source added to Communities. The directory stands at 925.
25 September 2026Directory
1 source added to Data. The directory stands at 924.
25 September 2026Site
The cabinet on the homepage is joined by a claw crane full of people and a rocket ride of trends, and the three together download in fewer bytes than the cabinet did alone.
24 September 2026Site
Every name in a search result now links to the source, on the All tab of Search and Resources as on every other tab. The communities on the homepage link out too, and each person there links the place they post, as on the People page.
17 September 2026Directory
1 source added to Institutions. The directory stands at 923.
17 September 2026Site
The site can be searched: a box on the homepage and a Search page in the header look through every source at once. Resources now holds the institutions and the press as well as the data, which puts a hundred sources for markets with no page of their own, and the global bodies that belong to no market, somewhere a reader can browse to for the first time.
17 September 2026Site
This change log has an All button, so the whole list is one press away after filtering it. The invitation to submit a source on the Media and People tabs is now a link, and the directory's entries for 11, 15 and 16 September, which an export fault had left out, are restored.
16 September 2026Directory
3 sources added to Data (2) and Institutions (1). The directory stands at 922. 1 conversation corrected.
15 September 2026Directory
2 sources added to Institutions. The directory stands at 919. 2 conversations corrected.
11 September 2026Directory
30 sources added to Communities (27) and Data (3). The directory stands at 917. 5 sources corrected.
10 September 2026Directory
1 source corrected.
9 September 2026Directory
1 source added to Institutions. The directory stands at 887.
7 September 2026Directory
12 sources added to Media (4), People (3), Press (3) and Communities (2). The directory stands at 886.
7 September 2026Site
The trend list over MCP pages its results, dates each row, and matches a capitalised word.
7 September 2026Site
The source search over MCP pages its results, fifty at a time.
7 September 2026Site
Method names the studio and the collaborators behind the research, with links.
7 September 2026Site
This change log. It replaces the list of what was added with everything that changed: rows removed and corrected as well as added, and the site's own changes beside them.
7 September 2026Site
A source can be submitted through a form rather than by email, from a button in the footer beside the download.
6 September 2026Directory
1 source added to Data. The directory stands at 874.
1 September 2026Site
A weekly scan of the directory's own sources looks for news that moves the tracked trends, and the first run's findings landed on the trends page.
1 September 2026Site
The trends page says when the trends were last compiled.
29 August 2026Site
A page explains how to connect an assistant over MCP, and every row count quoted in the prose is read from the export rather than typed.
28 August 2026Directory
6 conversations added. The tab now tracks 37.
28 August 2026Site
The newsletters have their own shelf on Media.
27 August 2026Directory
60 sources added to Institutions (39), People (16) and Media (5). The directory stands at 873.
27 August 2026Site
The directory is served over MCP, so an assistant can query it directly, and every row has a permanent name it can be cited by.
26 August 2026Site
Every book in the directory shows its cover.
24 August 2026Site
Every page is now written out at build time, so a crawler or an assistant that runs no JavaScript reads the same page a browser does.
24 August 2026Site
The Landscape, Media and People pages say what is in each drawer, in numbers counted from the data rather than remembered.
24 August 2026Site
About is written.
24 August 2026Site
Each tab on Media, People and Landscape has its own address, so a link to Books lands on Books.
24 August 2026Site
Corrections is written, and About is in the footer.
24 August 2026Site
Both tables show every row, rather than the first twenty-five and a button.
24 August 2026Site
Method and the download page say when the directory was last compiled, with the row counts beside the date.
24 August 2026Site
Market pages send a phone a phone-sized photograph rather than the desktop one.