JOB OPPORTUNITIES API — INTEGRATION BRIEF FOR AN LLM OR CODING AGENT Rendered from the live ledger on 2026-10-02. Always current at: https://api.jobopportunitiesapi.org/public/ai-prompt (text/plain, no key) Paste this whole block into your model or agent before asking it to write code against this API. It is written to be sufficient on its own: every endpoint, every parameter, every gate and every failure mode below has been sent against production. If anything here disagrees with the machine-readable contract at https://api.jobopportunitiesapi.org/v1/openapi.yaml — OpenAPI 3.1, no key needed to read it — the spec wins and this brief is the bug. Worked, tested example programs: https://github.com/lucagiftzek/joa-examples ================================================================================ 0. RULES FOR THE ASSISTANT USING THIS BRIEF ================================================================================ * Do not invent parameters, fields or endpoints. The complete lists are in sections 3, 5 and 6. An unrecognised query parameter is IGNORED, not rejected — so a hallucinated filter returns 200 with results that quietly do not match what the user asked for. That is the worst failure mode here and the only one the API cannot warn you about. * Do not suggest an endpoint the caller's plan cannot reach. See section 4. * Do not present an inferred value as a fact the employer stated. Every row carries `field_sources` (section 5) saying which is which. If the user asks "how many remote jobs", ask whether they mean employer-stated (`remote_confirmed=true`) or including our inference. The difference is 11.3% versus 89.8% of live rows. * Never estimate a salary. This API publishes only figures a source carried or that were read out of an advert's own text; it never models one. If a row has no salary the honest answer is that it has none, and section 8 says how many rows that is. * Never hard-code a count. Read `/public/stats`, `/public/coverage` or `/v1/meta/freshness`. Numbers in this brief are rendered from the ledger when you fetch it and are already stale by the time you paste it. * Poll `/v1/changes`, never `/v1/jobs`, to keep a copy in sync (section 7). * Records are metered per row returned (section 9). Ask for `limit` you need. ================================================================================ 1. BASE URL AND AUTHENTICATION ================================================================================ Base URL https://api.jobopportunitiesapi.org Auth Authorization: Bearer YOUR_API_SECRET Missing or bad key -> 401 {"error":"unauthorized"} curl -H "Authorization: Bearer $JOA_KEY" \ "https://api.jobopportunitiesapi.org/v1/jobs?country=DE&limit=5" Get a key: https://jobopportunitiesapi.org/register The free Explore plan is a real key against the live ledger — the same rows the paid tiers return, no card, no expiry, no trial clock. There is also a KEYLESS mirror under /public/ for the read-only endpoints in section 3. It is genuinely usable for evaluation and genuinely bounded: * one page per filter — the response carries no cursor, and passing one returns 402 * `limit` is clamped to 50 (a larger value is not an error; you get 50) * roughly 2 requests a second per address * `status=closed` and `status=any` are refused with 403 `paid_feature` Use /public/ to evaluate. Use /v1/ with a key to build. ================================================================================ 2. WHAT THE DATA IS, IN ONE PARAGRAPH ================================================================================ 2,875,004 live openings from 190,356 employers, collected from employer applicant-tracking systems (Greenhouse, Lever, Workday, Ashby, Personio, SmartRecruiters and ten more), employers' own careers pages, and public employment agencies. There is no aggregator or job-board inventory: that material is not redistributable, so it never enters the ledger. Every row carries where it came from, when it was last re-confirmed at that source, and which of its fields were stated versus derived. Rows that leave their source are kept, with a date and a reason, rather than deleted. ================================================================================ 3. ENDPOINTS — THE COMPLETE LIST ================================================================================ WITH A KEY (/v1) GET /v1/jobs a page of listings, newest first GET /v1/jobs/{id} one listing, including its description GET /v1/jobs/closed roles that left their source, newest closure first GET /v1/jobs/expired ids + closed_at + closed_reason only (cheap) [gated] GET /v1/companies employers with live roles GET /v1/companies/{slug} one employer GET /v1/changes delta feed: created/updated/withdrawn/delisted [gated] GET /v1/export full corpus as NDJSON [gated] GET /v1/meta/facets every filter value with its live count GET /v1/meta/providers every source, with live row counts GET /v1/meta/freshness how recently the ledger was verified GET /v1/me your key, plan, limits and usage today GET /v1/openapi.yaml the contract GET /v1/openapi.json the same, as JSON WITHOUT A KEY (/public) — same rows, bounded as described in section 1 GET /public/jobs GET /public/jobs/{id} GET /public/companies GET /public/companies/{slug} GET /public/facets GET /public/providers GET /public/freshness GET /public/coverage GET /public/coverage/countries GET /public/coverage/employers GET /public/stats GET /public/plans GET /public/status GET /public/ai-prompt (this document) GET /public/openapi.yaml GET /public/openapi.json ================================================================================ 4. PLAN GATING — DO NOT SUGGEST A 403 ================================================================================ Two entitlements decide everything: `delta_feed` and `bulk_export`. Read the caller's own from GET /v1/me before recommending an approach; read the plan matrix from GET /public/plans. plan price/mo records/mo req/day req/min delta_feed bulk_export explore free 1,000 5,000 30 no no growth EUR 80 60,000 100,000 120 yes no signal EUR 299 400,000 400,000 300 yes yes scale EUR 899 2,000,000 1,000,000 600 yes yes ON THE FREE EXPLORE PLAN THESE ARE 403 — do not build a design around them: GET /v1/changes -> 403 plan_upgrade_required (needs delta_feed) GET /v1/jobs/expired -> 403 plan_upgrade_required (needs delta_feed) GET /v1/export -> 403 plan_upgrade_required (needs bulk_export) The 403 body names the feature and the response carries the header `X-JOA-Required-Feature: delta_feed | bulk_export`, so a client can branch on it rather than on a parsed message. THESE DO WORK ON EXPLORE, and are the free-plan substitute for the delta feed: GET /v1/jobs/closed closure history, most recently closed first GET /v1/jobs?status=closed the same half of the ledger, in posting order GET /v1/jobs?status=any both halves, `status` on every row GET /v1/jobs?posted_after=... new rows since a timestamp GET /v1/jobs?verified_after=... rows re-confirmed since a timestamp `status=closed` and `status=any` are paid-only on the KEYLESS /public mirror (403 `paid_feature`), but available to every key including Explore. ================================================================================ 5. THE PROVENANCE OBJECT — THE POINT OF THIS API ================================================================================ Every job carries `field_sources`, one entry per field, with three values: "published" the source carried this value. An employer or agency wrote it. "inferred" we derived it. Ours, not theirs. "absent" there is none. Not zero, not unknown-therefore-false: none. "field_sources": { "remote": "published", "employment_type": "absent", "category": "inferred", "seniority": "inferred", "salary": "published", "location": "published", "posted_at": "published", "description": "published", "source_type": "inferred" } What that means in practice: * `category` and `category_confidence`, and `seniority`, are ALWAYS inferred. They are read off the job title by a classifier. Treat them as a search aid, never as an employer's own taxonomy. `require_fields=category` is refused with 422 for exactly this reason, rather than returning an empty page that would look like a coverage gap. * `remote` may be either. `remote_inferred: true` on a row, or `field_sources.remote == "inferred"`, means we worked it out from the listing text. `?remote_confirmed=true` narrows to employer-stated only. * `?require_fields=salary,description` returns ONLY rows where every named field is "published", and adds a `completeness` object to the response saying how many live rows carry each field across the whole ledger. Related row-level provenance, separate from field_sources: `source` the exact system, e.g. greenhouse, lever, workday, company_site `source_type` ats | career_site | public_agency `provider_type` the internal class behind source_type `last_verified_at` when we last confirmed the vacancy still exists at source `first_seen_at` when it entered the ledger `closed_at`, `closed_reason` expired_upstream (proved dead) | not_seen (stopped appearing) ================================================================================ 6. PARAMETERS OF GET /v1/jobs — THE COMPLETE LIST ================================================================================ Comma-separated means OR within the parameter; different parameters AND together. Values are matched case-insensitively. Timestamps are RFC3339 UTC. PAGING limit 1-200, default 25. Max 50 when include_description=true (51 is a 422, not a silent clamp). cursor the `next_cursor` from the previous page. See section 7. status live (default) | closed | any PLACE country ISO-3166 alpha-2, e.g. DE,FR,NL state two-letter US state codes, e.g. OH,TX. Absent where we could not establish it; ambiguous city names resolve to NULL on purpose, so this under-reports rather than misplaces a job. city e.g. Berlin remote remote | hybrid | on_site | not_stated remote_confirmed true = only rows whose remote status the SOURCE stated WHAT THE ROLE IS category 25 job families. `uncategorised` selects rows with no confident classification. seniority Entry | Mid | Senior | Lead | Manager | Director | Executive | Intern | not_stated employment_type Full-time | Part-time | Contract | Temporary | Internship | not_stated title full-text over the JOB TITLE only title_exclude drop rows whose title matches these words q full-text over title, company name and location — NOT the description. Stemming is off, so q=engineer does not match engineering. description_contains full-text over the advert BODY. Implies has_description. WHO IS HIRING company company slugs company_domain bare domains, e.g. stripe.com,figma.com — the join key you already have in a CRM. A scheme or path is tolerated. Employers without a known domain can never match. provider exact source system, e.g. greenhouse,lever,workday. Up to 12. An unpublished name is 422, never an empty page. source_type ats | career_site | public_agency (legacy spellings employer_ats | government | direct are still accepted) MONEY has_salary true (= structured, employer-published) | any (also figures read out of the advert text). AI estimates are never published under any value. min_salary lower bound on salary_min_annual_eur max_salary upper bound on the same field TIME posted_after date or RFC3339 verified_after only rows re-confirmed at source since this instant COMPLETENESS AND SIZE require_fields comma-separated; only rows where EVERY named field is "published". category and seniority are refused with 422. has_description true = only rows carrying an advert body include_description true = return the full text. Off by default: descriptions average ~2.5 KB, so a 200-row page would be ~500 KB. EXCLUSIONS (same vocabulary as their positive twin) exclude_category exclude_country exclude_provider exclude_source_type exclude_company_domain OTHER ENDPOINTS' PARAMETERS /v1/jobs/closed limit, cursor, closed_after, closed_before, closed_reason /v1/jobs/expired since (required), limit /v1/changes since (required), limit (1-5000, default 500) /v1/export after, status, and every filter above /v1/companies limit, offset, cursor, country, org_type, source_type, has_website, include_discovered, q ================================================================================ 7. PAGINATION — THREE DIFFERENT CURSORS, DO NOT MIX THEM ================================================================================ (a) /v1/jobs, /v1/jobs/closed, /v1/companies — OPAQUE CURSOR Keyset on (posted_at DESC NULLS LAST, id DESC). Send the response's `next_cursor` back as `cursor`. Stop when `has_more` is false. There is no offset on /v1/jobs and that is deliberate: at this scale an offset double-counts when rows shift under you. Treat the cursor as opaque; do not parse, construct or truncate it. cursor = None while True: r = GET /v1/jobs?country=DE&limit=100 [&cursor=cursor] yield from r["data"] if not r["has_more"]: break cursor = r["next_cursor"] (b) /v1/changes and /v1/jobs/expired — SINCE / NEXT_SINCE Pass `since` on the first call and the returned `next_since` from then on. It is a keyset cursor of "|", so no row is lost to a shared timestamp. An empty page echoes your cursor back, so the loop keeps working once you are current — that is how you know you are caught up, not an error. `change` is one of: created | updated | withdrawn | delisted. delisted the vacancy is gone. Stop serving it. withdrawn the vacancy may still exist, but we no longer stand behind the link — usually it turned out to point at a portal rather than the employer. Stop serving it too. The row keeps its id and can come back later as `updated`. IF YOUR CODE SWITCHES ON `change`, HANDLE withdrawn EXPLICITLY. An unknown-value branch that ignores it will keep serving rows we withdrew. (c) /v1/export — AFTER NDJSON, one object per line, ordered by `id`. Resume an interrupted transfer with `after=`. Records are billed as they are written, so check the LAST LINE of the stream: if the monthly allowance ran out mid-export it carries `error: record_quota_exhausted` and the `after` cursor to resume from. Do not assume a clean end of stream. ================================================================================ 8. HOW COMPLETE THE DATA IS — SAY THIS BEFORE A USER FINDS OUT ================================================================================ descriptions 88.9% of live rows (2,554,516) employer-stated salary 5.1% (146,422 rows) employer-stated remote 11.3% — the other 89.8% is our inference Salary is a narrow filter by nature, not a broken one. `salary_min_annual_eur` is one number in one unit, normalised from sources that quote USD an hour, GBP a year and EUR a month; it is null, never a guess, when the period or currency is unknown. `min_salary`/`max_salary` operate on it, so they select only rows that have it. The rates and the hours-per-year used are published in /v1/meta/freshness so the arithmetic can be checked. Per-country coverage, including where it is thin, is at /public/coverage/countries — keyless, so check it before you spend a record. ================================================================================ 9. LIMITS, METERING AND EVERY ERROR YOU CAN GET ================================================================================ Two independent budgets: REQUESTS per day/minute, and RECORDS per month. Records are charged per row actually returned, so `limit=200` costs 200 records and `limit=5` costs 5. Response headers on every keyed request: X-RateLimit-Limit requests allowed today X-RateLimit-Remaining requests left today X-RateLimit-Reset unix seconds until the day rolls over X-RateLimit-Records-Limit records allowed this month ("unlimited" is possible) X-RateLimit-Records-Used records charged so far this month X-RateLimit-Records-Remaining records left this month Retry-After seconds, on a 429 Errors are always {"error": "", "message": "", "docs": ...} 401 unauthorized missing or invalid key 402 record_quota_exhausted monthly RECORD allowance spent. Not a rate limit; waiting does not help. Also what /public returns if you try to page past the first page. 403 plan_upgrade_required your plan lacks delta_feed or bulk_export. Header X-JOA-Required-Feature names which. 403 paid_feature status=closed/any on the keyless /public mirror. 422 bad_provider a provider slug we do not publish 422 bad_source_type outside ats | career_site | public_agency 422 bad_category outside the `family` facet 422 bad_seniority outside the `seniority` facet 422 bad_remote outside remote | hybrid | on_site | not_stated 422 bad_employment_type outside the `employment` facet 422 bad_state not a US state code 422 bad_status status outside live | closed | any 422 field_never_published require_fields=category or seniority: both are inferred by construction, so no row can satisfy it 429 rate_limited requests per day or per minute exceeded A wrong vocabulary value is refused BY NAME rather than returning an empty page, because an empty page leaves you unable to tell "no jobs match" from "you spelled it wrong". Enumerate every legal value from /public/facets (keyless) or /v1/meta/facets, and every provider slug from /public/providers. ================================================================================ 10. PERFORMANCE — THE TWO THINGS THAT ARE SLOW ================================================================================ There is a 10-second statement timeout. Exceed it and you get 500 `query_failed`. It is deterministic, so DO NOT RETRY the identical request — it costs another ten seconds to reach the same answer. Lower `limit` instead. `limit` is what tips a query over, more than the filters are. Measured: ?description_contains=kubernetes&country=DE&limit=3 200 in 1.8 s ?description_contains=kubernetes&country=DE&limit=5 200 in 2.0 s ?description_contains=kubernetes&country=DE&limit=8 500 in 10.1 s ?description_contains=kubernetes&country=DE&limit=25 500 in 10.1 s The cost is per matching row FOUND, so a rare match is expensive to page through. For a narrow full-text query, ask for 5 rows at a time, not the default 25. * `title=` and `description_contains=` together is the most expensive pair this API can be asked for: both indexes are GIN and the matching rows must still be fetched to be ordered by posted_at. Send it narrowed by country or category, with a small `limit`, never against the whole ledger. * Do not send a filter that another filter already implies. It cannot remove a row, and the extra predicate is enough to flip the plan. * Deep pagination is fine — the cursor is a keyset, not an offset — but asking for descriptions in bulk is not. Use `include_description=true` with `limit<=50` on a dense query (a country plus a category), and a much smaller `limit` on a sparse one. ================================================================================ 11. FIVE RECIPES THAT WORK ================================================================================ New German engineering roles from the last day, employer-stated remote only: GET /v1/jobs?country=DE&category=Engineering&remote_confirmed=true&posted_after=2026-08-14&limit=50 Everything one employer has open, by the domain in your CRM: GET /v1/jobs?company_domain=stripe.com&limit=100 A salary benchmark with no estimates anywhere in it: GET /v1/jobs?require_fields=salary&country=DE&category=Engineering&limit=200 then read salary_min_annual_eur; the `completeness` block gives the denominator Keep a local copy in sync (needs delta_feed): first: GET /v1/changes?since=2026-08-01T00:00:00Z&limit=500 after: GET /v1/changes?since=&limit=500 apply created/updated as upserts, delisted/withdrawn as removals Which sources the data actually comes from, before you pay anything: GET /public/providers GET /public/coverage/countries ================================================================================ 12. TERMS ================================================================================ Listings are facts about vacancies and are returned for use in your product. Do not re-publish the corpus as a competing feed. Full terms: https://jobopportunitiesapi.org/terms — plans and prices: https://jobopportunitiesapi.org/pricing — a person reads support@jobopportunitiesapi.org.