https://your-site.com/blog/hello is, in your database,
post id 42.
That matters, because every update Castro sends goes to PUT /posts/{id}. Without
the id, Castro can see a page on your site but cannot touch it.
Page sync closes that gap. Castro asks your site for a list of its pages,
with ids, and stamps each entity’s identity onto the matching crawled URL.
What you implement
One endpoint, gated by thepages.list capability.
GET /pages
Return every published, publicly reachable entity you want Castro to be able
to update: posts, products, categories, authors. Support per_page and page
query parameters (Castro asks for 100 at a time).
string
required
Your id: the same one you returned from
POST /posts or POST /products.
This is the whole point of the endpoint.string
required
Used as a fallback match when the URL doesn’t line up exactly.
string
The absolute, canonical public URL. This is what Castro matches against its
crawl, so it must be the URL your site actually serves.
string
Stored against the page in Castro and shown throughout the UI.
string
Your own type label. Castro maps the well-known ones onto its page
classification. See the table below.
string[]
Optional. Category names. Omit it and sync still works; the field just stays
empty in Castro.
string[]
Optional, same as
categories.How type is interpreted
Sending a label Castro doesn’t recognise is not an error. It just means Castro
trusts what its crawler worked out instead.
What happens when sync runs
1
Castro pages through GET /pages
100 entities at a time, a few requests in parallel. Castro stops when it gets
a short page, so the last page must be shorter than
per_page: the
standard pagination contract. A site with a million pages is streamed, never
buffered.2
Each page is matched by URL
Against the URLs Castro crawled. This is the fast path, and it’s why the
url
you return must be the canonical one your site serves.3
Anything unmatched falls back to slug
Last path segment, compared case-insensitively: this rescues the usual
mismatches (trailing slash,
www, http vs https).4
Still unmatched pages get queued for a crawl
A published page Castro has never crawled isn’t dropped. Castro queues it,
crawls it, and the next sync matches it. Nothing is lost.
When sync runs
You don’t have to do anything to trigger it. It fires from three places:The Sync button
Settings → Integration → Custom Website → Sync pages from your site.
Only appears if you declared
pages.list.Right after connecting
So a freshly connected site is identified immediately.
After every crawl
New pages appear in the crawl, so their identity is stamped straight away.
Making sure Castro can find your pages
Sync stamps identity onto pages Castro has crawled. If Castro’s crawler never reached a page, sync has nothing to attach the id to (it’ll queue it, but that’s a slower round-trip). Help the crawler:- Publish a
sitemap.xmllisting every public page. - Make sure the URLs in your sitemap are the same ones you return from
GET /pages. - Don’t put page content behind JavaScript that only renders client-side, if you can avoid it.
Verify it
The conformance script checks thatGET /pages returns a bare JSON
array, that a published post appears in it, and that the post’s title survived a
partial update, which is what page sync reads back.
Once sync has run, open any crawled page in Castro. It should show your entity id
and title. If it shows neither, work through the
page sync troubleshooting.

