100% local — your data never leaves your browser

JSON to BigQuery Schema — bq load Takes It

Turn a JSON sample into a BigQuery schema file. Fields from every row are merged, so the load succeeds on the first attempt rather than the third.

Instant Private Zero cookies
Indentation

JSON input

BigQuery output

What this tool does

A JSON sample becomes a BigQuery schema file — the array of field descriptions that bq load --schema expects.

{ "id": 1, "firstName": "Ada", "address": { "city": "Paris" } }
[
  { "name": "id", "type": "INTEGER", "mode": "REQUIRED" },
  { "name": "firstName", "type": "STRING", "mode": "REQUIRED" },
  {
    "name": "address",
    "type": "RECORD",
    "mode": "REQUIRED",
    "fields": [{ "name": "city", "type": "STRING", "mode": "REQUIRED" }]
  }
]

Field names keep the JSON keys — BigQuery accepts letters, digits and underscores, which is what most keys already are.

Every row, not just the first

[{ "a": 1 }, { "a": 2, "b": "x" }]
[
  { "name": "a", "type": "INTEGER", "mode": "REQUIRED" },
  { "name": "b", "type": "STRING", "mode": "NULLABLE" }
]

The schema used to be read off the first object alone, so loading the rest of your own sample failed with no such field: b. Fields are now merged across every object, and one that is missing anywhere is NULLABLE — which is what it is.

That was a real defect, found while writing this page and fixed, along with the two below.

Types are merged along with the fields. A key typed INTEGER on one row and STRING on the next gets a JSON column — STRING would refuse the number at load time. Two objects, on the other hand, merge into a single RECORD, and a subfield missing from either side is NULLABLE there.

A mixed array is a JSON column

{ "arr": [1, "x"] }
[{ "name": "arr", "type": "JSON", "mode": "NULLABLE" }]

A repeated column has one type, and [1, "x"] has two. Typing it from the first element — INTEGER REPEATED — produced a schema that the sample itself does not load into. BigQuery’s JSON type holds the value as it stands, so that is what mixed arrays and arrays of arrays get. Integers mixed with decimals stay a repeated FLOAT, which covers both.

Two guesses, named

  • null becomes STRING NULLABLE. The sample shows a key and no type; STRING is a choice, not a reading. If the field is a number, change it before the first load.
  • An empty array becomes STRING REPEATED. Same reasoning: nothing was in it. Both load your sample fine, and both are worth a second look when you know the field.

Loading your JSON into BigQuery

Three things have to line up, and BigQuery names none of them clearly when they do not.

One object per line. BigQuery loads newline-delimited JSON, not a JSON array. If your file starts with [, it will not load — the NDJSON converter here does exactly that conversion.

The schema in its own file. Save the array above as schema.json. It is not part of the data file, and it is not read from it.

bq load \
  --source_format=NEWLINE_DELIMITED_JSON \
  --schema=schema.json \
  mydataset.mytable \
  data.ndjson

Or through the API. The same array goes under schema.fields of the load job configuration — the file and the API take identical content.

Private by design

Everything runs locally in your browser with JavaScript. Your data is never uploaded, which makes the tool safe for sensitive content, and it keeps working offline.

Frequently asked questions

Why did my mixed array become a JSON column?
Because no repeated column can hold it. `[1, "x"]` used to be typed `INTEGER REPEATED`, from the first element — and then the load of that very sample failed on the second. BigQuery has a native `JSON` type that takes the value as it stands, so that is what a mixed array gets. A homogeneous one still becomes a repeated column of its type.
Why is a field NULLABLE that my first row always has?
Because another row does not. The schema is the union of every object in the sample, and a field missing from some of them cannot be REQUIRED — the load would reject those rows. Reading only the first object gave a schema that failed on the rest with "no such field".
How do I use the file?
`bq load --source_format=NEWLINE_DELIMITED_JSON --schema=schema.json mydataset.mytable data.ndjson`. The schema lives in a file of its own, beside the data rather than inside it; through the API, the same array goes under `schema.fields`.

Related converters