Median

Connected sources

Import Notion pages, GitHub repositories, docs sites and websites, and keep them current.

Updated Oct 5, 202611 minute read

Open Knowledge, press Add, then Add content. The Sources list has Notion, GitHub, Docs site and Website.

Only admins and owners can open a source. Everyone else sees Admin access required. See roles.

Refresh and limits

SourceRefreshes on its ownBy handLimits
NotionWhen the agent starts a reply, at most every 15 minutes. Only pages already importedLibrary sync button20 pages per import. Sub-pages 3 levels deep, 25 per page
GitHub repositoryOn every push to the synced branch. When the agent starts a reply, at most once a minute per repositorySync now, library sync button500 markdown files. 900 KB per file
Docs siteSame as a repositorySync now, library sync button500 pages
WebsiteNeverRe-scrape500 pages per crawl. 3 crawls at once
  • A refresh runs beside the reply. The reply uses what is indexed at that moment, and the next reply gets the changes.
  • The library sync button has the tooltip Pull fresh copies from GitHub and Notion, or names just the one the library holds. It reads every file in every repository and docs site again, and pulls every Notion page again while Notion is connected.
  • Sync now also reads every file, not just the ones that changed. Pressed during a sync, either button runs once that sync finishes, and so does a push.
  • Titles come from a file's frontmatter title, then its first heading, then its filename. It appears once the library holds a GitHub document, or a Notion page with Notion still connected.
  • A repository or docs site whose last sync failed shows up under Needs fixing on the Dashboard as owner/repo is not syncing.

Notion

Connect

Open Notion and press Connect Notion. Notion's consent screen opens. Choose the pages Median may see there.

Pick pages

Under Pages, tick pages in the list, or type in Search pages, or paste a link. The list shows 20 pages at a time, most recently edited first. A pasted link imports that one page.

Import

Leave Include sub-pages on to bring child pages along. Press Import. The library opens and each page arrives with an Indexing badge.

RuleValue
Pages per import20. More returns "Up to 20 pages at a time."
Sub-page depth3 levels below each picked page
Sub-pages per page25
Sub-page ceilingNew sub-pages stop once the organization holds 300 Notion documents
TitleThe page's name in Notion
EditingEdit in Notion. The document page shows a Read only badge with the tooltip "Edit this document in Notion."

A refresh pulls every imported page again. It does not add new pages or new sub-pages. Import those.

Choose pages on the Workspace row reopens Notion's consent screen, where you can share more pages with Median.

Disconnect asks Disconnect Notion? and keeps every imported document. Refreshes stop.

Connect errorCause
Notion connection was not authorized. Try connecting again.Consent was declined on Notion
The connection did not finish. Try again.Notion did not complete the handoff

A connect link works once, for 15 minutes. After that, the return from Notion lands on a page that reads "This connection link has expired or was already used. Start again from Median."

Error on a pageCause
This page has nothing on it yet.The page is empty
This page is more than one document can hold. Split it up.The page converts to more than 900 KB of markdown
Notion is not connected anymore. Connect it, then try again.The workspace was disconnected
The Notion connection expired. Connect it again, then retry.Notion would not renew access
That does not look like a link to a Notion page.The pasted link is not a Notion page

A failed page shows Failed in the library. The reason and Try again are on the document page.

GitHub

Median reads repositories through its GitHub app. You pick which repositories the app can see during the install on GitHub.

Connect an account

Open GitHub and press Connect GitHub. Install the app on GitHub and pick repositories. You come back to the GitHub page.

Add a repository

Press Add repository, find it with Search repositories, and press Import.

Set the fields

Pick a Branch and a Folder, fill Published at if you want links, and press Add. The first sync starts at once.

FieldDefaultNotes
BranchThe repository's default branchThe default branch is followed even if it is renamed later
FolderSync the whole repositoryEach folder shows how many markdown and spec files it holds
Published atEmptyOptional. Where the files are published, so the agent can link them. See Published links
  • Connect another under Accounts adds a second GitHub account or organization. The account menu in the repository picker has Connect another account too.
  • The picker lists recently active repositories. Type a full owner/repo to reach one the list does not show.
  • Missing a repository? Add access on GitHub opens the app's settings on GitHub.
  • One repository can be added more than once with different folders. The same repository and folder twice returns "That repository is already connected."

What a sync takes

FileBecomes
.md, .mdx, .markdown under the folderOne document per file. The title is the first heading, or the file name
An OpenAPI or Swagger specOne document per endpoint, plus an overview. See API specs
  • A sync compares every file with the last sync. Unchanged files are skipped. Changed files are indexed again.
  • A file deleted from the repository leaves the library.
  • A new file joins the library folder its neighbours in the repository are in.
  • Markdown files over 900 KB are skipped.
  • Documents whose indexing failed are retried on every sync.
  • Synced documents are edited in the repository. The document page shows a Read only badge with the tooltip "Edit this document in its source repository."

Repository rows

Each row shows owner/repo, the folder and branch when set, and a status line.

StatusMeans
Syncing nowA sync is running, including the first one after you add it
Synced 5m agoThe last sync finished
The last sync failedThe error shows under the row
ButtonDoes
Published docsOpens the address dialog. See Published links
Sync nowStarts a sync. Does nothing while one is running
RemoveAsks Remove owner/repo?, then removes the repository and its documents

Disconnect on an account removes every repository and docs site read through it, and their documents. Uninstalling the app on GitHub, or taking a repository's access away there, does the same.

Sync errors

ErrorFix
More than 500 markdown files in there. Point the sync at a docs folder.Remove the repository and add it again with a folder
This docs site has more than 500 pages, which is more than a sync can take.The site grew past 500 pages. Remove it and add the repository as a plain repository with a smaller folder
This repository is too large to sync. Select a documentation folder instead.GitHub returned a partial file tree for the branch. The whole branch is read even when a folder is set, so a folder does not clear this
GitHub could not find that. Check the repository and branch, and that the app can see it.The branch is gone, or the app lost access
GitHub turned us away. Check the app's key, and that it is still installed.The app was uninstalled or lost permission
GitHub is not answering right now. Try again in a bit.GitHub was down. Press Sync now later
Could not sync this repository. Try again.Anything else

A sync that runs longer than 10 minutes is treated as stuck. The next sync takes over.

Connect errorCause
This connection attempt expired. Connect again.More than an hour passed on GitHub
That GitHub installation is connected to another organization. Disconnect it there first.One GitHub install belongs to one Median organization

API specs

Repositories and docs sites turn OpenAPI specs into documents. Each endpoint document has Parameters, Request body, Responses and a curl command under Example request. The overview document lists the version, Base URL, Authentication and every endpoint.

Each spec can also go on your site as an API reference.

A file is read as a spec when all three tests pass:

TestPasses
Extension.json, .yaml or .yml
Name or folderThe name contains openapi or swagger. Or the name has api as its own word, bounded by the start or end or by -, _ or ., so api.json and public-api.yaml pass and rapid.json does not. Or a parent folder is named openapi, swagger, spec, specs or apis
ContentsIt parses, has an openapi or swagger version key, and has a paths object

A file that fails the contents test is ignored.

LimitValue
Spec files read per sync25
Spec file size5 MB
Endpoints per spec300. The rest are left out
VersionsOpenAPI 3.x and Swagger 2.0

Specs must sit under the synced folder. A docs site also takes specs its config names, wherever they sit.

Docs sites

A docs site is a GitHub repository read through its framework's config. The config sets which files are pages and the address each page publishes at. It uses the same GitHub accounts as the GitHub source.

Open Docs site and press Add docs site. The dialog has three steps.

StepWhat you do
RepositoryPick the repository. Continue
FrameworkCheck the Branch. Median reads the branch and shows the Framework with its config file, and Pages with any API specs. Continue
AddressFill Published at. It is required. Three sample files show the address each will publish at, and update as you type. Add docs site

A Blume or Docusaurus config that names its site fills Published at for you.

FrameworkConfigPages and routes
Mintlifydocs.json or mint.jsonOnly pages in the navigation, at the route the navigation gives. Specs from any openapi field
Docusaurusdocusaurus.config.js, .ts, .mjs, .cjs or .mtsFiles under docs.path, default docs. Published under routeBasePath, default docs. Number prefixes like 01- are dropped. url plus baseUrl fill the address. Specs from specPath
Fumadocssource.config.ts, .js, .mjs or .mtsFiles under defineDocs({ dir }), default content/docs. Published under the folder holding the [[...slug]] route, default docs
GitBook.gitbook.yaml or SUMMARY.mdOnly pages the summary links. root and structure.summary are read from the yaml
Blumeblume.config.ts, .js, .mjs, .mts or .jsonFiles under content.root, default docs. deployment.site fills the address. Specs from openapi when enabled: true
  • index and README files publish at their folder's route, except on Mintlify.
  • A route prefix is not added twice when Published at already ends with it.
  • When a repository holds several configs, GitBook is tried last.
  • Median reads configs as text and does not run them. Check the three sample addresses before you add the site.
SituationWhat you see
No config foundNot recognised, "Median reads Mintlify, Docusaurus, Fumadocs, GitBook and Blume." and a Sync it as a plain repository link to the GitHub source
Over 500 pages"This site has more than 500 pages, which is more than a sync can take." Continue stays disabled. Add it as a plain repository with a smaller folder instead
Config names more than 100 API specs"This site exceeds the page limit for a single sync."

The site row shows the framework, the content folder and the branch, with the same status line and buttons as a repository row. Sync now on a docs site reads the config again even when the branch has not changed. Unchanged files are still skipped. Every sync that finds the repository changed reads the config again too, so a moved docs folder follows.

Websites

Websites need a paid plan. Without one the page shows "Available on Standard and Pro." See plans.

Paste an address

Open Website and type the address, for example docs.example.com. https:// is added when missing.

Scrape

Press Scrape. The crawl appears under Sites, and pages land in the library as they are scraped.

  • The crawl stays under the address you give. example.com/docs does not read example.com/blog.
  • It reads the site's sitemap and follows links. PDFs are read too.
  • Each scraped page is charged in credits. A crawl stops at 500 pages, or sooner when credits run short.
  • Each page becomes a document linked to the address it was scraped from.

A running crawl opens into three stages.

StageShows
DiscoverPages found
ScrapePages scraped, as 12 of 40
IndexPages indexed. Starts once scraping ends
CrawlRow reads
Running"Started 2m ago by" and the name of whoever started it
Finished"Scraped 5m ago by" and the same name. The page count, a skipped count for pages that came back blank or too large, and a kept count for pages you edited sit on the right
FailedThe error, with the same counts on the right
ButtonWhenDoes
CancelDiscover or Scrape runningStops the crawl. Pages already scraped stay and are indexed
DismissAfter scraping endsRemoves the row. The pages stay in the library
Re-scrapeCrawl finished or failedFetches live copies and updates pages in place. Pages you edited keep your version
  • Nothing refreshes on a schedule. Press Re-scrape when the site changes.
  • A first crawl may use copies up to 2 days old. A re-scrape, or a new crawl of an address you scraped before, fetches live pages.
  • The list shows the 20 most recent crawls.
ErrorCause
That does not look like a site address.Not a web address
That site is already being scraped.A crawl of that address is running
Up to 3 sites at a time. Let one finish first.3 crawls are running
Nothing readable came back from this site.No page had content
The crawl took too long, so it was let go. Try again.The crawl ran for 60 minutes
The page import service is busy. Try again in a minute.The scraper was busy
The scraped pages could not be indexed. Try the site again.Indexing failed after the scrape
Your organization is out of credits. Add more in Settings, Billing.No credits were left when the crawl started

Other failures say to try again later.

From the API

RequestDoes
POST /v1/knowledge/syncSame as the library sync button
POST /v1/integrations/reposAdds a repository with repo, branch, path and publishedAt
PATCH /v1/integrations/repos/{owner}/{repo}Sets Published at
DELETE /v1/integrations/repos/{owner}/{repo}Stops syncing a repository. Its documents stay in the library, unlike Remove in the dashboard
GET and POST /v1/knowledge/crawlsLists crawls or starts one with url
DELETE /v1/knowledge/crawls/{id}Cancels or dismisses a crawl

Full schemas are in the management API reference.

Still need help?

    Esc