Connecting a database or a file store read-only: tables an agent consults during a run, documents read into a run, results delivered back, and the same tables on your own machine.
A data source is a read-only connection from a workspace to a store the customer already has — the ERP's database, the SFTP folder the scanner drops invoices into, a bucket, a spreadsheet. It feeds a run in two ways, and a third piece sends the results back:
data source (read-only)
|-- saved table ------------.
'-- documents --------------+
(input.sourceObjects) |
v
the run
| save_output, report
v
connection (sftp, s3, email, webhook)
-> recorded as the run's deliveriesThere are ten kinds of external store, plus uploaded files:
postgres (PostgreSQL), mysql (MySQL / MariaDB), mssql (SQL Server). Read by SQL query, or a table or view by name.s3 (Amazon S3, or any S3-compatible endpoint), azure_blob (Azure Blob Storage), azure_files (Azure Files). A CSV, XLSX or JSON file reads as a table, and documents can be read into a run.ftp (FTP or FTPS), sftp (SFTP). As object storage: files read as tables, documents read into a run.gsheets (Google Sheets): a sheet tab reads as a table. bigquery (BigQuery): a SQL query, or a table by name.file (Uploaded files): the files you upload onto the source, read as tables or into a run as documents.Only the file stores (ftp, sftp, s3, azure_blob, azure_files, file) hand over documents; a database, BigQuery or a spreadsheet gives tables only. The group names and labels are the console's own (Data sources → Add a data source).
A file is parsed as CSV (delimiter sniffed, so TSV and plain-text exports work), XLSX/XLSM or JSON, by its content type first and its extension second.
Nothing on this page can change the customer's data. That is not a setting; it is the shape of the code:
SELECT or WITH; INSERT, UPDATE, DELETE, CREATE, SET, INTO and the rest are rejected even inside a CTE, and so are row locks (FOR UPDATE, FOR SHARE). A rejected query is 400 read_only_violation and never opens a connection.BEGIN TRANSACTION READ ONLY with a statement_timeout; MySQL/MariaDB sessions are switched to SET SESSION TRANSACTION READ ONLY before any query, or the read fails. SQL Server has no read-only transaction mode, so connect it with a login mapped to db_datareader. Google Sheets and BigQuery authenticate with the spreadsheets.readonly and bigquery.readonly scopes.directory; anything that would climb out of it (.., an absolute path) is refused, not rewritten.config and sealed at rest (AES-256-GCM) before the row is written. No response ever returns them — only their field names, in secretFieldsPresent — no audit row records them, and no model sees them. A driver error that could echo a host or a connection string reaches you as 502 upstream_unreachable with a fixed message.data_source.read audit row with the query text or the object id — never a credential. Per workspace, reads are limited to 30 a minute and connection tests to 10 a minute.s3 source's custom endpoint must resolve to a public address.In the console: Data sources → Add a data source. Four steps — pick the kind of source, enter its connection details, test it, done. Create & test creates the source and runs one read-only connection test against it; until the test passes, the source stays a draft you can correct and test again.
Over the API:
curl -X POST "https://api.agentflowbind.com/v1/data-sources" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" -H "Content-Type: application/json" \
-d '{
"type": "postgres",
"name": "ERP replica",
"config": {
"host": "erp.example.com",
"port": 5432,
"database": "acme_erp",
"user": "afb_reader",
"password": "...",
"sslMode": "require"
}
}'{
"id": "4c2888cd-...",
"name": "ERP replica",
"type": "postgres",
"status": "active",
"projectId": "0f6b1c2a-...",
"metadata": { "host": "erp.example.com", "port": 5432, "database": "acme_erp", "user": "afb_reader", "sslMode": "require" },
"secretFieldsPresent": ["password"],
"lastTestAt": null,
"lastTestOk": null,
"lastUsedAt": null,
"createdAt": "2026-09-17T09:00:00.000Z",
"updatedAt": "2026-09-17T09:00:00.000Z"
}name is unique in the workspace (409 otherwise). The config fields per type — the ones marked * are secret, sealed as above:
postgres — host, database, user, password*; optional port (5432) and sslMode: disable, require (the default) or verify-full.mysql — host, database, user, password*; optional port (3306) and sslMode: disable or require (the default).mssql — host, database, and a SQL Server login's user and password*; optional port (1433), encrypt, trustServerCertificate.s3 — region, bucket, accessKeyId*, secretAccessKey*; optional endpoint (blank means AWS — set it for MinIO, R2, Wasabi…), sessionToken*, prefix.azure_blob — accountName, container, and the one secret authMethod names: sasToken* (the default), accountKey* or connectionString*; optional prefix.azure_files — accountName, shareName, and the one secret authMethod names: sasToken* (the default) or accountKey*; optional directory.ftp — host, user, password*; optional port (21), protocol (ftp or ftps), directory, passive.sftp — host, user, and password* or privateKey*; optional port (22), passphrase* (with a key), directory.gsheets — serviceAccountJson*, spreadsheetId (the id or the sheet's URL); optional sheetName, range.bigquery — serviceAccountJson*, projectId; optional dataset, location.file — nothing; upload files to it (below).curl -X POST "https://api.agentflowbind.com/v1/data-sources/4c2888cd-.../test" -H "Authorization: Bearer afb_live_xxxxxxxxxxxx"{ "ok": true, "latencyMs": 41 }A source that can't be reached is still a 200 — { "ok": false, "latencyMs": ..., "error": "..." } — because a failed probe is the answer, not an API error. The outcome is kept on the source (lastTestAt, lastTestOk).
PATCH /v1/data-sources/:id takes name, config (merged onto what is stored; a secret field you leave out keeps its sealed value, null clears it), projectId and status. { "status": "disabled" } pauses a source: a run cannot take its documents or tables, and an agent reading one of its tables gets a warning instead. DELETE /v1/data-sources/:id returns 204.
A file source has no connection; you upload the files it serves — CSV, XLSX or JSON to read as tables, or documents to read into a run — 25 MB each, virus-scanned before they are stored:
curl -X POST "https://api.agentflowbind.com/v1/data-sources/9e03.../files" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" \
-H "Content-Type: text/csv" -H "X-Filename: price-list.csv" \
--data-binary @price-list.csv# What the source holds — tables and views, or files and folders under a prefix
curl "https://api.agentflowbind.com/v1/data-sources/4c2888cd-.../objects?prefix=public" -H "Authorization: Bearer afb_live_xxxxxxxxxxxx"
# One object's columns
curl "https://api.agentflowbind.com/v1/data-sources/4c2888cd-.../schema?object=public.purchase_orders" -H "Authorization: Bearer afb_live_xxxxxxxxxxxx"
# One read: a query (SQL sources) or an objectId (any source) — exactly one of the two
curl -X POST "https://api.agentflowbind.com/v1/data-sources/4c2888cd-.../read" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" -H "Content-Type: application/json" \
-d '{ "query": "SELECT order_no, supplier, net_amount FROM purchase_orders WHERE open_amount > 0", "limit": 500 }'{
"columns": [{ "name": "order_no", "type": "text" }, { "name": "supplier", "type": "text" }, { "name": "net_amount", "type": "numeric" }],
"rows": [["OA-2026-0311", "Rossi Componenti", "13453.44"], "..."],
"rowCount": 13,
"truncated": false,
"summary": { "rowCount": 13, "columns": ["..."], "sampleRows": ["..."] }
}Only postgres, mysql, mssql and bigquery accept a query; the others read by objectId. limit is clamped to the row cap; truncated and notice say when a cap was reached. The errors specific to reads:
400 read_only_violation — the query is not a single SELECT/WITH read.400 unsupported_object — the object exists but can't be turned into a table.408 query_timeout — the source did not answer within the time cap.502 upstream_unreachable — the source could not be reached (the message never carries its config).A saved table is a read you name once and replay whenever it is needed: a query on a SQL source, or a picked object on any other. It stores the definition and a snapshot of the columns and row count — never the rows: every use reads the source again, under the same caps and the same read-only guard as the read above. In the console, open a source, run a query or pick an object, then Save as table; the source's Tables tab lists them.
curl -X POST "https://api.agentflowbind.com/v1/datasets" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" -H "Content-Type: application/json" \
-d '{
"dataSourceId": "4c2888cd-...",
"name": "Open purchase orders",
"query": "SELECT order_no, supplier, net_amount FROM purchase_orders WHERE open_amount > 0"
}'objectId instead of query saves a picked object (stored as definition.objectKey). A table belongs to its source's project and moves with it; its name is unique in the workspace. POST /v1/datasets/:id/materialize replays the read now and answers with rowCount, columns, truncated, readAt and lastMaterializedAt — never the rows.
An agent whose config lists saved tables in datasetIds gets the read_dataset tool (an agent with none never sees it). In the console this is the agent's Data sources section: saved tables this agent may read while it runs. Set it when creating the agent (POST /v1/agents) or in a new version (POST /v1/agents/:id/versions stores the whole config you send — send the current config with the ids added); each id must be a table of the workspace the agent can reach.
During the run, read_dataset called with no argument lists the linked tables from their snapshots (no source read); called with a table's name or id, it replays that table's saved read and returns the rows. The model never writes a query — it can only name a table somebody saved. A table whose source is paused, broken or unreachable answers { "ok": false, "warning": "..." } and the run carries on with what it has; every read is audited. A run reads only the tables its own project reaches.
input.datasetId on POST /v1/runs reads the table before the run exists and attaches its rows as a CSV input, the same way an upload is attached. The run records input.dataset: { datasetId, rowCount, readAt }; a paused or unreachable source is an error on this request, before anything is queued or charged. In the console's run picker this is the Saved tables tab.
The documents a run processes can come straight off a file store — ftp, sftp, s3, azure_blob, azure_files or file:
curl -X POST "https://api.agentflowbind.com/v1/runs" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" -H "Content-Type: application/json" \
-d '{
"agentId": "3c1a9e40-...-77bd",
"input": {
"sourceObjects": [
{ "dataSourceId": "9b1f7c30-...", "objectId": "inbox/fattura-rossi-componenti-2026-0184.pdf" },
{ "dataSourceId": "9b1f7c30-...", "objectId": "inbox/fattura-hansen-2026-0417.pdf" }
]
}
}'objectId is an id from GET /v1/data-sources/:id/objects. Every object is read before the run row is created, so anything wrong is an error on this request rather than a failed run:
POST /v1/files, checked on the content, not on the name), and one the agent's preset accepts;Accepted documents become ordinary input files. The run keeps where each came from in input.sourceObjects — dataSourceId, objectId, filename, sizeBytes, readAt — and each read writes a data_source.read audit row. The object itself is left exactly where it was: the inbox is read, never emptied.
The errors, all before the run exists:
404 source_object_not_found — the object is not on that source (any more).404 not_found — no such source for this caller.400 invalid_request with details.code source_object_unsupported — the source is a database, BigQuery or Google Sheets: save a table and attach that instead.400 invalid_request with details.code source_object_too_large — the document is over 25 MB.400 invalid_request with details.code input_mime_rejected — the platform, or this agent, does not accept the file's type.In the console: New run → Input → From a data source — choose the source, browse its folders, tick documents, or pick a folder whole: its direct files (one level) come in, up to 25.
Writing a result next to the documents is a delivery, and a delivery goes through a connection — never through the data source, which stays read-only. For the SFTP inbox above, the matching connection is an sftp one pointing at the output folder:
curl -X POST "https://api.agentflowbind.com/v1/connections" \
-H "Authorization: Bearer afb_live_xxxxxxxxxxxx" -H "Content-Type: application/json" \
-d '{
"provider": "sftp",
"name": "Invoices · processed",
"config": { "host": "sftp.example.com", "user": "afb", "privateKey": "-----BEGIN OPENSSH PRIVATE KEY-----\n...", "directory": "/fatture/processed" }
}'Then name it on the run: "input": { ..., "fields": { "saveTarget": "<connection id>" } }. The agent's save_output step delivers its file (JSON, CSV or XLSX) there during the run; with "outputConfig": { "method": "report" }, the rendered report PDF follows to the same folder once the run has succeeded. GET /v1/runs/:id lists every delivery in deliveries — delivered with its ref (sftp://afb@sftp.example.com:22/fatture/processed/...), or failed with the reason — and the run page shows them under Delivered to. The report's delivery is best-effort: if it fails, the run stays succeeded and the failure is listed with the others.
afb-runner runs the same agent on a folder of the customer's own machine. The documents go only to the model provider configured there, under the customer's own key — or nowhere, with a model server on their own network — and Agent Flow Bind receives counts (files, duration, tokens per model), the seat and the licence state: never a file name, a path or a value. An agent's linked saved tables are resolved locally. The agent definition the runner downloads (GET /v1/agents/:id/definition, with the runner:run scope) carries each linked table's saved read and its source's type and name — never the cloud connection, its host or its credentials.
Two environment variables, both optional:
AFB_RUNNER_DATABASE_URL resolves every table saved as a query, through the same adapters as the platform — so the same read-only guard, read-only transaction, row cap and timeout. A postgres://, postgresql:// or mysql:// URL, for example postgres://afb_reader:...@10.0.0.4:5432/erp.AFB_RUNNER_TABLES_DIR resolves every table saved as an object, matched by file name in that folder: suppliers.csv, price-list.xlsx, payments.json (a sheet as price-list.xlsx#Prices).export AFB_RUNNER_DATABASE_URL='postgres://afb_reader:...@10.0.0.4:5432/erp'
export AFB_RUNNER_TABLES_DIR=/srv/tables
afb-runner doctor --agent <agent-id> --in /srv/invoices # checks first
afb-runner run --agent <agent-id> --in /srv/invoices --out /srv/results # a folder -> JSON + CSV + report
afb-runner ui # the same, from a local page?sslmode=require (or verify-full, PostgreSQL only) when the server needs it. Point it at a read-only user anyway.doctor tests the local connection (Connessione dati locale) and the tables folder (Cartella tabelle locali), and lists the linked tables this machine cannot serve (Tabelle collegate). A bare DATABASE_URL in the environment is a hard failure: it would redirect the runner's own store — use AFB_RUNNER_DATABASE_URL.afb-runner ui serves one page on 127.0.0.1 only (port 8765 by default), behind a random token in the address.agents:read and runner:run.To install it, open the console's Local runs tab → Install the runner: it shows the image to pull and the command line for each of your agents.
The console's own words for everything above, in its three languages:
| English | Italiano | Deutsch |
|---|---|---|
| Data sources | Origini dati | Datenquellen |
| Add a data source | Aggiungi un'origine dati | Datenquelle hinzufügen |
| Create & test | Crea e prova | Anlegen & testen |
| Test connection | Prova la connessione | Verbindung testen |
| Save as table | Salva come tabella | Als Tabelle speichern |
| Saved tables | Tabelle salvate | Gespeicherte Tabellen |
| From a data source | Da un'origine dati | Aus einer Datenquelle |
| Settings → Connections | Impostazioni → Connessioni | Einstellungen → Verbindungen |
| SFTP folder | Cartella SFTP | SFTP-Ordner |
| Delivered to | Consegnato in | Zugestellt an |
| Local runs | Esecuzioni locali | Lokale Ausführungen |
| Install the runner | Installa l'esecutore | Ausführer installieren |
The scopes an API key needs:
data_sources:read — GET /v1/data-sources and /v1/data-sources/:id, GET .../objects, GET .../schema, POST .../read; GET /v1/datasets and /v1/datasets/:id, POST /v1/datasets/:id/materialize.data_sources:manage — POST /v1/data-sources, PATCH and DELETE /v1/data-sources/:id, POST .../test, POST .../files; POST /v1/datasets, DELETE /v1/datasets/:id.runs:write — POST /v1/runs with input.sourceObjects or input.datasetId.webhooks:manage — /v1/connections (see Connections).In the console, managing sources, saved tables and connections takes a workspace owner or admin, or a project's admin or editor; anyone who can see a project can browse its sources and read its tables. Every request and response shape is in the API reference: data sources, saved tables, connections, runs.
The SDK (preview) wraps the data-source routes as client.dataSources — list, create, get, update, delete, test, listObjects, getSchema, read:
// SDK (preview)
const { dataSources } = await client.dataSources.list();
const erp = dataSources.find((source) => source.name === "ERP replica")!;
const probe = await client.dataSources.test(erp.id);
const orders = await client.dataSources.read(erp.id, { objectId: "public.purchase_orders" });Saved tables, connections and input.sourceObjects have no typed SDK method yet; call those routes directly.