Watch → Source → Chunk → Embed → Vector with no Clean is fine. Node names and descriptions are free-text, so labels like “Watch new docs” work.
The Trigger node decides when a refinery runs. Schedule fires at set times, Watch reacts when a bound source changes, Webhook and API run when an external system calls in, and the canvas [Run] button runs it on demand.
![Foundry canvas showing connected Schedule, Cloud Source, Redact, Chunk, Embed, and Vector nodes, with [Add node] and [Publish] buttons in the toolbar](https://mintcdn.com/conscience-technology/40NFQYSYq0qo9Q4J/images/docs/foundry-pipeline.png?fit=max&auto=format&n=40NFQYSYq0qo9Q4J&q=85&s=3f332c51c244b4cd2604dd367876b7a6)
When to reach for Foundry
- You have a folder of PDFs, spreadsheets, or web pages you want the Agent to answer from.
- You want Google Drive, Microsoft 365, or Slack content kept in sync with your knowledge.
- You want database rows tagged and made retrievable.
- You want to strip names, emails, phone numbers, and other PII before anything hits knowledge.
The six node kinds
Skip nodes that don’t apply. If wiring is off, a red or yellow badge appears in the top-right of the canvas so you can fix it before publishing.
Refineries vs Flows
A Flow runs an Agent. A refinery prepares the data the Agent uses. You don’t build agent behaviour in Foundry, and you don’t build data ingest in a Flow. The split keeps each surface focused on one job.Publishing a refinery
Same as Flow. Edit the Draft, click [Publish] and that snapshot becomes the Published version. Scheduled and webhook runs use the Published version, so editing the canvas doesn’t disturb runs already in flight. Rollback works the same as Flow.What to do next
Triggers
Learn how Schedule, Watch, Webhook, and API start refineries.
Connectors
Google Drive, Microsoft 365, Slack, Discord.
Normalize node
Break documents into retrievable chunks.
Run logs
Inspect a run and stored reports.