Databases / DuckDB

Natural language to SQL for DuckDB

NLQueries is an open source natural language to SQL engine for DuckDB. Ask a question in plain English, get SQL that is validated against your real DuckDB schema before it runs, and see the answer. It works from the command line, as a Python library, and as an MCP server for Claude Desktop, Cursor and other AI assistants.

Get started View on GitHub

Connect DuckDB and ask your first question

Install the package, register the connection once, build the knowledge base, and query. Passwords are prompted for interactively and stored in your OS keychain, never in a config file.

$ pip install "nlqueries-core[duckdb]"
$ nlqueries connect duckdb --database /data/warehouse.db --alias duck-local
$ nlqueries process-history duck-local --annotate && nlqueries export-kb duck-local
$ nlqueries query duck-local "how many orders were placed last month by first-time customers?"

How NL2SQL works on DuckDB

Most text-to-SQL tools hand the LLM a schema dump and hope. NLQueries builds a YAML knowledge base from your DuckDB schema and from real query history, in this case none; DuckDB is file-based and keeps no query history across connections, so it learns which joins your team actually uses and what your columns mean. That knowledge base is what the model sees when it writes SQL.

Every generated statement is parsed and checked column-by-column against the live schema with sqlglot before execution, then run inside a read-only transaction with a statement timeout. Use nlqueries ask to see the SQL without executing it, or nlqueries query to run it. A semantic cache answers repeated questions without another LLM or database round-trip.

DuckDB-specific details

There is no host, port or password: --database is a file path, and omitting it gives you a transient in-memory database.

Because there is no persisted query history, the knowledge base comes from schema introspection: primary keys via duckdb_constraints(), with foreign keys skipped since analytics workloads rarely declare them. Add relationships to the YAML knowledge base by hand for best results.

Full connector notes are in the connectors guide, and the read-only role to create is in database hardening. This connector covers DuckDB files and in-memory DuckDB databases.

Use it from Claude, Cursor or any MCP client

Run nlqueries mcp-server and point Claude Desktop, Cursor, or any Model Context Protocol client at it. The assistant can then query DuckDB through the same validated pipeline, with OIDC authentication and per-tool authorization available for network transports. See MCP authentication.

Frequently asked questions

How does NLQueries turn a natural language question into SQL?

It retrieves the relevant tables, columns, relationships and past query patterns from a YAML knowledge base built from your schema and query history, asks an LLM for SQL with that context, validates the SQL against your real schema with sqlglot before it runs, and executes it inside a read-only transaction with a statement timeout.

Is the generated SQL safe to run against production?

Generated SQL is validated column-by-column against the live schema, executed read-only with a timeout, and can be previewed with nlqueries ask without touching the database. Pair that with a least-privilege database role and it is designed for production use.

Can I query Parquet or CSV files with natural language through DuckDB?

Yes, if they are exposed as tables or views in the DuckDB file you connect. Create the views in DuckDB first, then run nlqueries export-kb so the schema is in the knowledge base.

Can Claude Desktop or Cursor query this database through NLQueries?

Yes. nlqueries mcp-server exposes the same pipeline as a Model Context Protocol server, so any MCP client can ask questions and get validated SQL results as a native tool call.

Other engines: Postgres · MySQL · Snowflake · BigQuery · Redshift · SQL Server · SQLite