What this tool does
A JSON sample becomes a BigQuery schema file — the array of field descriptions that bq load --schema expects.
{ "id": 1, "firstName": "Ada", "address": { "city": "Paris" } }
[
{ "name": "id", "type": "INTEGER", "mode": "REQUIRED" },
{ "name": "firstName", "type": "STRING", "mode": "REQUIRED" },
{
"name": "address",
"type": "RECORD",
"mode": "REQUIRED",
"fields": [{ "name": "city", "type": "STRING", "mode": "REQUIRED" }]
}
]
Field names keep the JSON keys — BigQuery accepts letters, digits and underscores, which is what most keys already are.
Every row, not just the first
[{ "a": 1 }, { "a": 2, "b": "x" }]
[
{ "name": "a", "type": "INTEGER", "mode": "REQUIRED" },
{ "name": "b", "type": "STRING", "mode": "NULLABLE" }
]
The schema used to be read off the first object alone, so loading the rest of your own sample failed with no such field: b. Fields are now merged across every object, and one that is missing anywhere is NULLABLE — which is what it is.
That was a real defect, found while writing this page and fixed, along with the two below.
Types are merged along with the fields. A key typed INTEGER on one row and STRING on the next gets a JSON column — STRING would refuse the number at load time. Two objects, on the other hand, merge into a single RECORD, and a subfield missing from either side is NULLABLE there.
A mixed array is a JSON column
{ "arr": [1, "x"] }
[{ "name": "arr", "type": "JSON", "mode": "NULLABLE" }]
A repeated column has one type, and [1, "x"] has two. Typing it from the first element — INTEGER REPEATED — produced a schema that the sample itself does not load into. BigQuery’s JSON type holds the value as it stands, so that is what mixed arrays and arrays of arrays get. Integers mixed with decimals stay a repeated FLOAT, which covers both.
Two guesses, named
nullbecomesSTRING NULLABLE. The sample shows a key and no type; STRING is a choice, not a reading. If the field is a number, change it before the first load.- An empty array becomes
STRING REPEATED. Same reasoning: nothing was in it. Both load your sample fine, and both are worth a second look when you know the field.
Loading your JSON into BigQuery
Three things have to line up, and BigQuery names none of them clearly when they do not.
One object per line. BigQuery loads newline-delimited JSON, not a JSON array. If your file starts with [, it will not load — the NDJSON converter here does exactly that conversion.
The schema in its own file. Save the array above as schema.json. It is not part of the data file, and it is not read from it.
bq load \
--source_format=NEWLINE_DELIMITED_JSON \
--schema=schema.json \
mydataset.mytable \
data.ndjson
Or through the API. The same array goes under schema.fields of the load job configuration — the file and the API take identical content.
Private by design
Everything runs locally in your browser with JavaScript. Your data is never uploaded, which makes the tool safe for sensitive content, and it keeps working offline.