Everything, zipped
Every table on this page, as CSV and as JSON, plus a README that repeats all of this offline and the licence. If you are not sure which of the files below you want, you want this one.
Everything behind this site, in files you can open or in a live connection your assistant can query. Take the lot in one zip, take the spreadsheet, take the single table you actually need, or take none of them and point Claude at the endpoint. The licence is CC BY 4.0, which is the lawyerly way of saying go on then.
Directory last updated 6 October 2026 · 978 sources · 12 markets · 43 trends
Every table on this page, as CSV and as JSON, plus a README that repeats all of this offline and the licence. If you are not sure which of the files below you want, you want this one.
The same tables as a single workbook, a sheet each, top row frozen and filters already on. Opens in Excel, Numbers, LibreOffice and Google Sheets. For anyone who would rather sort than script.
No file at all. The whole directory is served live over MCP, free, with no key and no account: add one URL to Claude and it can search the sources, pull a market profile and read the trends while you talk to it. Six tools, and every answer arrives carrying its date, its link and the things it is not safe to conclude.
The directory itself. Every outlet, institution, dataset, podcast, book, person and community, with what it lets you find out, who owns it, what it costs, which markets it covers, and what to watch for. This is the table most people come for.
The twelve market profiles as one file: every sourced claim under the section it belongs to, with the link it came from, and the player population figure for each market carrying the definition that says who was counted. Nested rather than flattened, because a population figure without its definition is a number waiting to be misused.
What is moving the industry, each with what it is, why it matters, the markets it lands in and how far along it is. The further reading is its own file, one row per link, so you can sort a reading list instead of unpicking a cell.
Storefronts, consoles, livestreaming and UGC platforms, with who owns each one. Genres with their definitions, their subgenres, and the terms actually used in Chinese, Korean and Japanese. The example games are their own file, one row per game.
Ways of finding things that I wish someone had told me in month one. Each one says what to type, gives a query you can paste, and explains what it surfaces and why.
Every value each field can hold and how many rows carry it. Read this first if you are building something on top of the data: it is the fastest way to see that Japan has 139 rows and Italy has 24 and to decide whether that gap matters for what you are doing.
One row is one thing: a source, a claim, a trend, a game. Where a row carries a list of other things, that list gets a file of its own rather than a cell with five links glued into it. That is why the further reading on a trend and the example games in a genre are separate tables, each keyed by the name of the thing they belong to.
A list of plain words stays in its cell. Markets and languages are written into one cell with a semicolon and a space between them, because splitting three tokens into a table of their own helps nobody.
CSV is UTF-8 and RFC 4180, quoted where a value needs it. Plenty of these rows carry Japanese, Korean, Chinese and Arabic, so if your spreadsheet opens one and shows you a fence of question marks, it has guessed the encoding. The workbook above is the quicker fix.
JSON is the same data with the nesting left in. There is nothing in the JSON that is not in the CSVs; it just has not been flattened, which is what you want if you are feeding this to code or to a model.
An empty cell means the field was not collected for that kind of row, not that the answer is nothing. Cost is asked of data platforms, where the price decides whether you can use the thing at all. ISBN is asked of books. Neither is missing from a subreddit.
There is no ID column. Rows are identified by name, and the row order is export order, which is not a ranking of anything by anyone. If you need keys that survive the next pass, build them yourself from the name and the link.
Not checked
Every source here was opened and read before its description was written. What the source then claims is not checked. A figure you reach through this data carries exactly the authority of whoever published it, and moving it into a spreadsheet does not upgrade that. Where an owner has a stake in what gets reported, the row says so, and that is a flag to read carefully rather than a correction to the number.
Derived from The verification standard set out on Method, restated here because a file travels away from the page that explains it.
Two ways in, and they are good at different things. Connect the directory over MCP and ask, and the model reads the rows it needs at the moment it needs them, with no file to attach and nothing to go stale on your desk. Attach the file, and you have a fixed copy you can count, check, cite and rerun next month against the same rows. Reach for the connection when you are asking; reach for the file when you are working.
Whichever you use, do not paste the data into the chat. A model handed sources.csv as an attachment can filter it, count it and quote it. A model handed several hundred rows of it as text starts summarising them, and summarising is where invented rows come from.
The prompts below all do the same two things: they name the file and the columns, and they say what the model must not do. That second half is the half that matters, and it matters over the connection too. Left alone a model will rank these sources, fill an empty cell with something plausible, or quietly refresh a 2026 figure from memory, and nothing in the data gives it a way to notice. Take one, change the market or the question, and go. If you are connected rather than attached, drop the sentence about the file and keep everything after it.
I have attached files from Delightful's Game Research Starter Pack, a directory of where to research the video games industry. The sources file holds 978 rows, one per source, under the columns name, group, medium, format, what_it_lets_you_find_out, markets, language, cost and link. It is called sources.csv, and if it reached you under some other name, those columns are how you know you have the right file. Treat it as a map of where to look, not as a store of facts. Every source was read before it was described, but nothing the sources themselves claim has been checked, and nothing in the file is ranked or scored, so row order means nothing. When you answer, name the rows you used and give me the link from the row. If the file does not cover what I asked, say so instead of filling the gap from what you already know. Before I ask you anything, tell me what you have: how many rows the file holds, what values the group column takes and how many rows sit under each, and how many distinct market codes appear in the markets column, bearing in mind that one cell may hold several separated by semicolons. If you cannot work those out from the file itself, say so plainly, because that tells me it did not reach you in a form you can count and everything after it would be guesswork.
From sources.csv, give me everything covering Japan: the rows whose markets column contains JP. Group them by the group column, and inside each group put the rows whose language is Japanese first. Return a table of name, what it lets you find out, and link. Then tell me which languages I will need, and flag any group holding fewer than five rows, because that is thin coverage in this directory rather than a fact about Japan.
I need to state how much was spent on mobile games in Germany last year, and I need something citable. Using sources.csv, list every row that could plausibly publish that figure, starting with the Data and Institutions groups. For each one, tell me what the ownership_note column says where it is filled. Cost was only collected on the Data group, so it is empty on every Institutions row: where it is empty, say the file does not record it rather than reading that as free. Then order them by how likely I am to get the number without a subscription, and say where you are ordering on an empty cell. Do not give me the figure itself. I am going to go and read it.
From sources.csv, take only the rows where group is Data and cost is Free or Freemium. Those are the two values that mean I can get something without paying, and if any row in that group carries no cost at all, list those separately as unknown rather than dropping them or ruling them in. For each, tell me what it measures, what it explicitly excludes, and which of them draw on the same underlying data so I do not check one number twice and think I have confirmed it. The what_it_lets_you_find_out column states the known limits. Use those rather than your own impression of these tools.
Using trends.csv and trend-reading.csv, joined on the trend name, write a one-page brief on the trends whose stage is Dominant or Emerging and whose markets include Korea or Japan. For each one, quote the figures in what_it_is exactly as written, add the why_it_matters line, and pick the two most useful links from the reading file, naming their type. Do not add a trend that is not in the file, and do not update any figure from your own knowledge. The as_of column tells you how fresh each row is; say so where it is old.
From genres.csv, pull the local_terms column for shooters, role-playing games and puzzle games, and give me the Chinese, Korean and Japanese terms with a romanisation. Then, using the how and example columns of searching-tips.csv, turn each term into two or three queries I can paste straight into a search engine to find market coverage rather than store pages. Keep the local script in the queries.
Using vocabulary.csv, tell me how many rows this directory holds per market and per language. Then name the markets where the count is low enough that an absence in this data says more about the directory than about the market itself. I am about to present findings from these files and I want to know where I would be arguing from a gap.
Creative Commons Attribution 4.0. Use it, cut it up, publish from it, build something you charge for. The one condition is credit: name this site, link the licence, and say if you changed something.
The licence covers the data in these files. It does not cover what the rows point at. The reporting, figures, artwork and writing at the far end of every link belong to whoever made them, and nothing here hands you any part of that.
Where the files came from and what was done to them. How the research ran, what Claude was and was not allowed to decide, what was verified, what deliberately was not, and the two problems in this data that no amount of care fixes.
Something wrong, something missing, something that should not be here at all? Tell me. Corrections go into the workbook and come out in the next pass, so the files change with the pages.
Every accepted export and every change to the site leaves a line here, newest first. The dates are the workbook's and git's, the counts are the ones the export recorded at the time, and none of it is typed by hand, so this list cannot quietly disagree with the files above it.
A directory line says what arrived, what was removed by name, and how many rows already here were corrected. A site line says what the site itself does differently. Show one kind or read both. For the story behind a correction, read Corrections.