Databases / BigQuery
Natural language to SQL for BigQuery
NLQueries is an open source natural language to SQL engine for BigQuery. Ask a question in plain English, get SQL that is validated against your real BigQuery schema before it runs, and see the answer. It works from the command line, as a Python library, and as an MCP server for Claude Desktop, Cursor and other AI assistants.
Connect BigQuery and ask your first question
Install the package, register the connection once, build the knowledge base, and query. Passwords are prompted for interactively and stored in your OS keychain, never in a config file.
$ pip install "nlqueries-core[bigquery]" $ nlqueries connect bigquery --project-id my-gcp-project --dataset-id my_dataset --service-account-json /path/to/service-account.json --alias bq-prod $ nlqueries process-history bq-prod --annotate && nlqueries export-kb bq-prod $ nlqueries query bq-prod "how many orders were placed last month by first-time customers?"
How NL2SQL works on BigQuery
Most text-to-SQL tools hand the LLM a schema dump and hope. NLQueries builds a YAML knowledge base from your BigQuery schema and from real query history, in this case INFORMATION_SCHEMA.JOBS, so it learns which joins your team actually uses and what your columns mean. That knowledge base is what the model sees when it writes SQL.
Every generated statement is parsed and checked column-by-column against the live schema with sqlglot before execution, then run inside a read-only transaction with a statement timeout. Use nlqueries ask to see the SQL without executing it, or nlqueries query to run it. A semantic cache answers repeated questions without another LLM or database round-trip.
BigQuery-specific details
There is no password. Omit --service-account-json and the connector uses Application Default Credentials, which is the natural fit on GCE, Cloud Run or a workstation that has run gcloud auth application-default login.
BigQuery has no transaction to roll back, so a non-SELECT statement type is logged after the job has run rather than prevented. Give the service account the BigQuery Data Viewer and Job User roles and nothing more.
process-history --days N is honoured because BigQuery tracks job execution time.
Full connector notes are in the connectors guide, and the read-only role to create is in database hardening. This connector covers Google BigQuery.
Use it from Claude, Cursor or any MCP client
Run nlqueries mcp-server and point Claude Desktop, Cursor, or any Model Context Protocol client at it. The assistant can then query BigQuery through the same validated pipeline, with OIDC authentication and per-tool authorization available for network transports. See MCP authentication.
Frequently asked questions
How does NLQueries turn a natural language question into SQL?
It retrieves the relevant tables, columns, relationships and past query patterns from a YAML knowledge base built from your schema and query history, asks an LLM for SQL with that context, validates the SQL against your real schema with sqlglot before it runs, and executes it inside a read-only transaction with a statement timeout.
Is the generated SQL safe to run against production?
Generated SQL is validated column-by-column against the live schema, executed read-only with a timeout, and can be previewed with nlqueries ask without touching the database. Pair that with a least-privilege database role and it is designed for production use.
Which IAM roles does the NLQueries service account need for BigQuery?
BigQuery Data Viewer on the dataset and BigQuery Job User on the project are enough to introspect the schema, read INFORMATION_SCHEMA.JOBS for query history, and run generated SELECT statements.
Can Claude Desktop or Cursor query this database through NLQueries?
Yes. nlqueries mcp-server exposes the same pipeline as a Model Context Protocol server, so any MCP client can ask questions and get validated SQL results as a native tool call.
Other engines: Postgres · MySQL · Snowflake · Redshift · SQL Server · DuckDB · SQLite