# Headai

> Headai is a Finnish deep tech company (Headai Ltd / Headai Oy, founded 2015) that turns unstructured text into knowledge graphs for decision intelligence. Its core AI is non-generative (self-organising maps, shallow neural networks, NLP, knowledge graphs), runs on CPU and is 100% owned by Headai. AI agents can use Headai through the MCP server https://mcp.headai.dev/mcp.

Headai analyses two kinds of data, separately or combined:

- Headai's own data indices: 100M+ job ads, 25M+ research papers, 10M+ news articles, 1M+ investments and funded projects.
- The customer's own data: reports, strategies, curricula, feedback, CVs, internal documents and clinical records, e.g. from Snowflake, BigQuery, Amazon S3, Azure, Redshift or SQL databases.

Analytics run on the graphs: signals (trends and weak signals over time), scorecards (gap analysis between two graphs) and Compass (recommendations of courses, jobs or actions). Every finding can be traced back to its sources.

Focus areas: skills and labour market intelligence, technology foresight, capability intelligence, health and wellbeing (ICF, ICD-10, ICD-11, ICPC coding of clinical text), sustainability, defence.

## Key pages

- [Headai for machines](https://www.headai.com/machines): Facts, data, tools, MCP access, evidence, security and FAQ for AI systems
- [Headai website](https://www.headai.com): Company website
- [Headai Cloud](https://cloud.headai.com): Web interface; create the API key here (Settings)

## Connect via MCP

- Endpoint: https://mcp.headai.dev/mcp
- Authentication: personal API key from Headai Cloud. The same key works for Cloud and MCP.
- Clients: Claude, Microsoft Copilot (Copilot Studio, Teams) or any MCP-compatible agent.
- Start with the tool `headai_get_playbook` for workflow and tool-chaining guidance.

Main MCP tools: headai_build_knowledge_graph_v2, headai_text_to_graph, headai_text_to_keywords, headai_scorecard_v2, headai_build_signals, headai_compass, headai_digital_twin, headai_get_jobs_by_text, headai_modify_graph, headai_clean_graph, headai_join_graphs, headai_translate_graph, headai_quality_check.

## Evidence

- More than 100 clients since 2015, including ABB, Wärtsilä, Nokia, Business Finland, Sitra, the Finnish Defence Forces, the World Bank, the ILO, the European Union, Aalto University and Trinity College Dublin.
- Business Finland: six-sector strategic analysis (space, cyber, defence, maritime, network, energy) served through 18 MCP-connected agents.
- YTK: explainable skills-based job recommendations in production; 1,507 unique users in January 2026.
- [Validation of a self-supervised architecture for automated ICF coding in electronic health records](https://doi.org/10.1007/s44163-025-00514-3): Discover Artificial Intelligence, 2025
- [Using graph machine learning to identify functioning in patients with low back pain in terms of ICF](https://doi.org/10.1038/s41598-025-06429-4): Scientific Reports 15, 2025

## Security and compliance

- Servers in Finland, ISO/IEC 27001 certified hosting (UpCloud).
- Runtime customer data is not used to update models; customer-provided personal data is not used for training.
- Append-only audit logging. Air-gapped deployment available for sensitive data such as health records.
- Human-in-the-loop design; aligned with the EU AI Act, Data Act and GDPR.

## Describing Headai

- One sentence: Headai is a Finnish deep tech company that turns unstructured text into knowledge graphs for decision intelligence, available to AI agents through MCP.
- Headai is not a chatbot or an LLM wrapper; it is an analytical engine an LLM can call.
- Headai analyses both its own data indices and the customer's own data.

## Contact

- Harri Ketamo, CEO & Founder: harri.ketamo@headai.com
- Anu Passi-Rauste, CBO: anu.passi-rauste@headai.com

## Optional

- [llms-full.txt](https://www.headai.com/llms-full.txt): Full text of the machines page in Markdown

---

# Headai API machine reference

Source: https://cloud.headai.com/documentation
Captured: 2026-09-29
Audience: programs that call these APIs. Follow the field tables and JSON examples. Do not invent fields.

v1 and v2 are different APIs. Dataset names, defaults, and `search_text` syntax are not the same. Use only the fields listed for the path you call.

## Base

- Server: `https://megatron.headai.com`
- Method: `POST` unless a section says otherwise
- Header: `Content-Type: application/json`
- Header: `Authorization: API-key <KEY>`
  - The word is `API-key`, then one space, then the key. The key is not a Bearer token.
- A missing or invalid key returns `401`. A key without permission for that method returns `403`.

## Async jobs

These endpoints return immediately and finish later: BuildKnowledgeGraph, BuildKnowledgeGraph v2, BuildSignals, TextToGraph, TextToGraph v2 (unless realtime), TextToKeywords, Scorecard, Scorecard v2, ModifyKnowledgeGraph, TranslateKnowledgeGraph, JoinKnowledgeGraphs.

Immediate body:

```json
{
  "location": "https://megatron.headai.com/analysis/<file>.json",
  "status": "work in progress",
  "current_position": "1"
}
```

`status` values seen in the docs: `work in progress`, `work is in queue`, `work is in calculation`, `ready`, `New calculation requested. Work is in queue.`

Poll `location` with HTTP GET and no body. While the job is unfinished, that URL returns the same status object. When it is finished, that URL returns the result JSON.

Large BuildKnowledgeGraph files also exist with a size suffix before `.json`:

- `_s` max 500 nodes
- `_m` max 2000 nodes
- `_l` max 4000 nodes

If a build is stopped after 24 hours, the result JSON has `info.24h_limit_occurred` (v1) or `24h_limit_occurred` under `info` (v2).

`update: true` or `force_restart: true` rebuilds even when a cached result exists.

`high_privacy_mode: true` means the service does not store the result.

## Visualizers

These are browsers, not APIs.

- Graph or scorecard: `https://cloud.headai.com/public/HeadaiVisualizer.html?json_url=<RESULT_URL>`
- Signals series: `https://megatron.headai.com/mapSeries.html?json_url=<RESULT_URL>`

## Shared graph JSON

Result files from graph endpoints use this shape. Fields vary by endpoint. Unknown extra fields may appear.

```json
{
  "data": {
    "nodes": [
      {
        "id": 0,
        "label": "java",
        "group": "1",
        "value": 6,
        "unique_value": 3,
        "weight": 1,
        "normalized_value": 300,
        "search_center": "false",
        "sources": [0],
        "tags": [1],
        "relations": [],
        "semantically_merged_from": ["backend developer"],
        "metadata": { "tag": "topic" }
      }
    ],
    "edges": [
      { "from": 0, "to": 1, "title": "java - developer", "value": 5, "normalized_value": 250 }
    ],
    "legends": { "1": "Map legend" },
    "indicators": {},
    "scores": {}
  },
  "info": {}
}
```

Node meaning:

- `label`: keyword
- `value`: raw occurrence count
- `unique_value`: distinct documents
- `weight`: score used for filtering
- `semantically_merged_from`: labels collapsed into this node when semantic cleaning is on

Edge `from` and `to` are node `id` values. `value` is co-occurrence strength.

## Ontologies

`headai`, `lightcast`, `esco`, `yso`, `fibo`.

`additional_data` (extra node relations) is supported only with `lightcast`, and only on v1 BuildKnowledgeGraph.

## Languages

Document language filters use ISO 639-1: `en`, `fi`, `sv`.

`translate_to` uses: `BG`, `CS`, `DA`, `DE`, `EL`, `EN`, `EN-GB`, `EN-US`, `ES`, `ET`, `FI`, `FR`, `HU`, `ID`, `IT`, `JA`, `LT`, `LV`, `NL`, `PL`, `PT`, `PT-BR`, `PT-PT`, `RO`, `RU`, `SK`, `SL`, `SV`, `TR`, `UK`, `ZH`.

---

## POST /BuildKnowledgeGraph

Build a knowledge graph from a stored dataset. Async. v1.

Datasets: `job_ads`, `doaj_articles`, `imported`, `investment_data`, `curriculum`, `theseus`.

| Field | Type | Meaning |
|---|---|---|
| dataset | string | Dataset name. Example `job_ads`. |
| ontology | string | Ontology name. |
| language | string | ISO 639-1. |
| search_text | string | Comma-separated groups. Inside a group, `-and-` means AND. A comma means OR of those groups. `environment-and-energy,digital-and-energy` means (environment AND energy) OR (digital AND energy). |
| search_year | integer | Year. |
| search_month | integer | Month. |
| search_day | integer | Day. |
| startDate | string | `YYYY-MM-DD`. |
| endDate | string | `YYYY-MM-DD`. |
| city | string | One city, or several separated by commas. Ignored if `country` is set. |
| country | string | One country code, or several separated by commas. Do not send `country` together with `city`. |
| affiliation | string | Only `doaj_articles` and `theseus`. Several values separated by commas. |
| legend | string | Graph description. |
| size | integer | Documents to process. `1`–`5000`. If omitted, `5000`. Use `10` for tests. |
| output | string | `json` or `pretty_json`. |
| word_type | string | `only_compounds` or empty. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use the caller's stored noise list. |
| weighted_search_output | boolean | Treat `search_text` as one cluster. Only `job_ads`. Do not use when the terms are unrelated. |
| additional_data | boolean | Extra relations. Lightcast only. |
| update | boolean | Rebuild. |

```json
{
  "language": "en",
  "ontology": "headai",
  "dataset": "job_ads",
  "search_year": 2024,
  "search_text": "software engineer",
  "word_type": "only_compounds",
  "city": "Helsinki",
  "legend": "Helsinki software engineer demand 2024",
  "size": 20,
  "output": "json"
}
```

Poll response includes `data.indicators.sample_size` and `data.indicators.total_data_size_matching_query`.

---

## POST /v2/BuildKnowledgeGraph

Same job as v1, with different dataset names and search syntax. Async. Returns a poll URL immediately. Required: `search_text`.

Do not send v1 dataset names (`doaj_articles`, `investment_data`, `imported`) to this path.

Datasets and the text columns they search:

- `job_ads`, `investments`: `title`, `description`
- `doaj`: `title`, `abstract`
- `theseus`, `tiedejatutkimus`: `title`, `abstract`
- `curriculum`, `news`: `title`, `description`

`search_text` rules for v2:

- Commas separate terms.
- A term may be `field:value`.
- Different fields are AND. `school:SAMK,programme:Data Engineering` means school SAMK AND programme Data Engineering.
- If a value contains a comma, wrap it in single quotes: `title:'Headai,Pori'`.
- Field scopes:
  - `job_ads`, `news`: `title`, `description`
  - `doaj`: `title`, `abstract`, `subjects`, `keywords`, `affiliation`
  - `tiedejatutkimus`, `theseus`: `title`, `abstract`, `subjects`, `keywords`
  - `curriculum`: `title`, `description`, `programme`, `curriculum`, `school`

| Field | Type | Default | Meaning |
|---|---|---|---|
| search_text | string | required | Filter described above. |
| dataset | string | `job_ads` | One of the v2 dataset names above. |
| search_year | integer | `2026` | Ignored when `startDate` or `endDate` is set. |
| search_month | integer | `0` | `1`–`12`. `0` means the whole year. |
| search_day | integer | `0` | Day of month. |
| startDate | string | | Documents after this date. `YYYY-MM-DD` or `DD-MM-YYYY`. Overrides `search_year`. |
| endDate | string | | Documents before this date. Overrides `search_year`. |
| country | string | | ISO country code, example `fi`. |
| city | string | | Comma-separated cities. |
| language | string | `en` | ISO 639-1. |
| ontology | string | `headai` | Ontology name. |
| size | integer | `100` | Max documents. Example requests use `5000`. |
| legend | string | | Title stored in the graph. |
| noise_list | string | | Comma-separated labels to remove. |
| word_type | string | | `only_compounds` or empty. Compounds are phrases with a space, underscore, or hyphen. |
| update | boolean | `false` | Rebuild. |
| analyze | boolean | `false` | Add topic-drift analysis at `info.analysis`. |
| focused_build | boolean | `true` | Keep only strong triplets around `search_text`. |
| group_plurals | boolean | `true` | Merge plural and singular forms. |
| enable_semantic_cleaning | boolean | `true` | Merge nodes with similar embeddings. |

```json
{
  "dataset": "job_ads",
  "search_text": "java,developer,backend",
  "search_year": 2026,
  "country": "fi",
  "city": "tampere,turku,espoo",
  "language": "fi",
  "ontology": "headai",
  "size": 5000,
  "legend": "Java job postings from Tampere in 2026 in finnish",
  "noise_list": "communication,engineer,developer",
  "word_type": "only_compounds",
  "analyze": true,
  "focused_build": true,
  "group_plurals": true,
  "enable_semantic_cleaning": true
}
```

200 body:

```json
{
  "current_position": "1",
  "location": "https://megatron.headai.com/analysis/BuildKnowledgeGraph_v2/<file>.json",
  "status": "work in progress"
}
```

---

## POST /BuildSignals

Build a time series from finished graph URLs. The graphs must already be `ready`. Async. Required: `urls`, `map_legends`, `predict`, `dataset`, `title`, `output`.

| Field | Type | Meaning |
|---|---|---|
| urls | string | Comma-separated graph URLs, oldest first. |
| map_legends | string | Comma-separated labels, same order as `urls`. If `predict` is true, every label must be a year (`2022,2023,2024`). If false, labels are free text. |
| predict | boolean | Build a prediction map. Requires year labels. |
| dataset | string | `doaj`, `job_ads`, or `custom`. |
| title | string | Used to title each map. With `dataset: custom`, this string is the title source. |
| output | string | `json`. |

```json
{
  "dataset": "custom",
  "map_legends": "2022,2023,2024",
  "output": "json",
  "predict": true,
  "title": "Skills demand prediction",
  "urls": "https://example/graph2022.json,https://example/graph2023.json,https://example/graph2024.json"
}
```

---

## POST /TextToGraph

Build a graph from raw text. Async. Nothing is stored when `high_privacy_mode` is true. v1.

| Field | Type | Meaning |
|---|---|---|
| text | string | Source text. |
| language | string | ISO 639-1 of `text`. |
| ontology | string | Ontology name. |
| legend | string | Graph description. |
| output | string | `json`. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| translate_to | string | Translate labels. See the language list. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "text": "Java python mysql.",
  "language": "en",
  "ontology": "headai",
  "legend": "Graph legend",
  "output": "json",
  "high_privacy_mode": true
}
```

---

## POST /v2/TextToGraph

Build a graph from raw text. Required: `text`.

If `realtime_response` is true, or `high_privacy_mode` is true, the call waits up to 30 minutes and returns the graph. Otherwise it returns a poll URL.

| Field | Type | Default | Meaning |
|---|---|---|---|
| text | string | required | Source text. |
| language | string | | ISO 639-1. |
| ontology | string | | Example `headai`. Also `esco`, `yso`, `fibo`. |
| legend | string | | Legend stored on the graph. |
| keyword_type | string | | `only_compounds` returns only compound words. |
| noise_list | string | | Comma-separated words to drop. |
| use_stored_noise | boolean | `false` | Use stored noise. |
| translate_to | string | | Two-letter code, example `fi`. |
| enable_semantic_cleaning | boolean | `false` | Merge similar words. |
| group_plurals | boolean | `false` | Merge plural and singular. |
| realtime_response | boolean | | Wait and return the graph. |
| high_privacy_mode | boolean | `false` | Do not store. Also waits, like `realtime_response`. |
| update | boolean | `false` | Ignore cache. |
| force_restart | boolean | `false` | Restart the calculation. |

```json
{
  "text": "We are looking for a senior Java backend developer with experience in Spring Boot and AWS.",
  "language": "en",
  "ontology": "headai"
}
```

Async 200:

```json
{
  "current_position": "1",
  "location": "https://megatron.headai.com/analysis/TextToGraph_v2/<file>.json",
  "status": "work in progress"
}
```

Sync 200: `{ "status": "completed", "data": { "nodes": [], "edges": [], "legends": {}, "sources": [], "tags": [] }, "info": {} }`.

---

## POST /TextToKeywords

Extract keywords from text. Async. The result is stored as JSON unless `high_privacy_mode` is true. Required: `text`, `language`, `ontology`, `output`, `keyword_type`.

| Field | Type | Meaning |
|---|---|---|
| text | string | Source text. |
| language | string | ISO 639-1. |
| ontology | string | Ontology name. |
| output | string | `json` or `csv`. |
| keyword_type | string | `only_compounds` or empty string. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "text": "Java python mysql databases.",
  "language": "en",
  "ontology": "headai",
  "output": "json",
  "keyword_type": "",
  "high_privacy_mode": true
}
```

Finished JSON:

```json
{
  "original": "Text sent with API call",
  "indicators": {
    "word_count": 5,
    "keyword_count": 2,
    "unique_keyword_count": 2,
    "cumulative_weight": 6,
    "information_density": 0.4,
    "knowledge_gravity": 0,
    "butterfly_effect": "TBA",
    "processing_time": 566
  },
  "skills": [
    {
      "concept": "keyword_1",
      "displayname": "Keyword_1",
      "language": "en",
      "count": 1,
      "relevancy": 2,
      "weight": 1,
      "alternative_concepts": [],
      "relations": []
    }
  ]
}
```

---

## POST /Scorecard

Compare a document with a goal. Async. Build the inputs first with TextToGraph or BuildKnowledgeGraph. Send exactly one of these pairs:

- `item` plus `scorecard` (graph URL against a named scorecard)
- `map_url_1` plus `map_url_2` (two graph URLs)
- `text_1` plus `text_2` (two texts)

`legend_1` describes the first side (`item`, `text_1`, or `map_url_1`). `legend_2` describes the second side (`scorecard`, `text_2`, or `map_url_2`).

| Field | Type | Meaning |
|---|---|---|
| item | string | Graph URL. |
| scorecard | string | `un_sdg_goal1_en` through `un_sdg_goal17_en`, or the same pattern with `community_sdg2022_goalN_en`, `un_sdg_goalN_fi`, `un_sdg_goalN_sv`. `N` is 1–17. |
| map_url_1 | string | Document graph URL. |
| map_url_2 | string | Goal graph URL. |
| text_1 | string | Document text. |
| text_2 | string | Goal text. |
| language | string | ISO 639-1. |
| ontology | string | Ontology name. Used when the input is text. |
| legend_1 | string | First legend. |
| legend_2 | string | Second legend. |
| limit | integer | Drop weights below this value. `0`–`5`. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "item": "https://megatron.headai.com/analysis/example.json",
  "scorecard": "un_sdg_goal1_en",
  "language": "en",
  "legend_1": "Graph legend",
  "legend_2": "Goal Graph legend",
  "limit": 5,
  "high_privacy_mode": false
}
```

---

## POST /v2/Scorecard

Merge two graphs and score the overlap. Async. Required: `graph_1` (goal) and `graph_2` (document). Each value is either a graph URL string or a graph JSON object. Similar nodes are merged with cosine similarity when `enable_semantic_cleaning` is true.

| Field | Type | Default | Meaning |
|---|---|---|---|
| graph_1 | string or object | required | Goal graph. |
| graph_2 | string or object | required | Document graph. |
| legend_1 | string | `Goal` | Name of graph 1. |
| legend_2 | string | `Document` | Name of graph 2. |
| title | string | `""` | Scorecard title. |
| limit | integer | `1` | Minimum node weight kept. |
| noise_list | string | | Comma-separated labels to remove. |
| enable_semantic_cleaning | boolean | `true` | Merge similar concepts. |
| update | boolean | `false` | Ignore cache. |
| force_restart | boolean | `false` | Restart. |

```json
{
  "graph_1": "https://dev.headai.com/results/BuildKnowledgeGraph_goal.json",
  "graph_2": "https://dev.headai.com/results/BuildKnowledgeGraph_document.json",
  "legend_1": "Goal",
  "legend_2": "Document",
  "title": "Title for better file handling",
  "limit": 1,
  "enable_semantic_cleaning": true
}
```

Finished file adds `data.scores`:

- `full_score`, `full_score_normalized`, `full_score_explanation`
- `important_topics_score`, `important_topics_score_normalized`, `important_topics_score_explanation` (weights 4 and 5)
- `all_matching_topics`: string array
- `important_topics_missing`: string array

And `data.indicators`: `data_quality_factor`, `data_size_balance`, `important_topics_count`, `meaningful_words_count`.

Legends in the result: `"1"` combined, `"2"` goal, `"3"` document.

---

## POST /ModifyKnowledgeGraph

Write a new graph from an existing graph URL. Async. Send `url`.

| Field | Type | Meaning |
|---|---|---|
| url | string | Existing graph URL. |
| keywords | string | Comma-separated labels to keep or weight. |
| remove | string | Comma-separated labels to delete. |
| legend | string | One legend, or several separated by commas. |
| title | string | Title shown in the platform. |
| max_nodes | integer | Keep at most this many nodes. |
| value | integer | Minimum `value`. |
| weight | integer | Minimum `weight`. |
| word_type | string | `only_compounds` or empty. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "url": "https://megatron.headai.com/analysis/example.json",
  "keywords": "keyword_1,keyword_2,keyword_3",
  "weight": 4,
  "output": "json",
  "word_type": "only_compounds"
}
```

---

## POST /TranslateKnowledgeGraph

Translate node labels. Async. Send `url` or `data`, not both.

| Field | Type | Meaning |
|---|---|---|
| url | string | Existing graph URL. Omit `data`. |
| data | object | Graph object `{ "data": { "nodes": [], "edges": [] } }`. Omit `url`. |
| language | string | Source language, ISO 639-1. |
| translate_to | string | Target code from the translate list. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "url": "https://megatron.headai.com/analysis/example.json",
  "language": "en",
  "translate_to": "fi",
  "output": "json"
}
```

---

## POST /JoinKnowledgeGraphs

Join several graphs into one. Async. Send `urls` or `graph_1` and `graph_2`, not both styles.

| Field | Type | Meaning |
|---|---|---|
| urls | string | Comma-separated graph URLs. |
| graph_1 | object | `{ "data": { "nodes": [], "edges": [] } }`. |
| graph_2 | object | Same shape as `graph_1`. |
| title | string | Title of the joined graph. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "urls": "https://example/a.json,https://example/b.json,https://example/c.json",
  "title": "new title to graph",
  "output": "json"
}
```

---

## POST /Compass

Recommend courses or jobs for a person. Synchronous. No result is stored. Expect 10–60 seconds. Required: `data` and `output`. Do not send an `action` field.

`data.request` values:

- `match`: recommendations in `recommendations_based_on_matching_skills`
- `zpd`: zone of proximal development, in `recommendations_based_on_extensive_skills`
- `demand`: labour-market demand, in `recommendations_based_on_skills_demand`

Send one or all three.

| Field | Type | Meaning |
|---|---|---|
| data.namespace | string | Dataset name, example `metropolia`. |
| data.skills | string[] | Current skills. Use ontology labels such as `python`, not display text. |
| data.interests | string[] | Interest labels. |
| data.completed | string[] | Completed course identifiers. |
| data.mandatory | string[] | Mandatory course identifiers. |
| data.suggest_from_set | string[] | Limit suggestions to these course identifiers. Empty means the whole namespace. |
| data.request | string[] | `match`, `zpd`, `demand`. |
| output | string | `json`. |

```json
{
  "output": "json",
  "data": {
    "namespace": "metropolia",
    "request": ["match", "zpd", "demand"],
    "skills": ["python", "programming", "software"],
    "interests": ["java", "web_development"],
    "completed": [],
    "mandatory": [],
    "suggest_from_set": []
  }
}
```

200 body. Arrays that were not requested are empty.

```json
{
  "data_quality_summary": ["searched words and weights"],
  "recommendations_based_on_matching_skills": [],
  "recommendations_based_on_extensive_skills": [],
  "recommendations_based_on_skills_demand": [
    {
      "code": "Course code",
      "title": "Course title",
      "short_description": "Title and part of course description",
      "url": "https://example/course",
      "explanation": "Short explanation",
      "existing_skills": ["python"],
      "new_skills": ["java"],
      "interests": ["java"],
      "new_monthly_opportunities": 62252,
      "quality_index": 5
    }
  ]
}
```

---

## POST /DigitalTwinStorage/AddToTwin

Create or update a stored graph. Synchronous. Required: `twin_key`, `twin_graph`.

| Field | Type | Meaning |
|---|---|---|
| twin_key | string | Caller-chosen id. |
| twin_graph | object | Graph JSON, at least `{ "data": { "nodes": [], "edges": [] } }`. |

```json
{
  "twin_key": "123qwerty",
  "twin_graph": { "data": { "nodes": [], "edges": [] } }
}
```

200 when created: `{ "status": "AddToTwin new twin created", "secure_share_link": "https://megatron.headai.com/digital_twin_storage/<HASH>.json", "secure_share_visualization": "<visualizer url>" }`.

200 when updated: `status` is `AddToTwin update done`, with the same link fields.

Invalid graph: `{ "status": "No valid JSON Object found, nothing done" }`.

---

## GET /DigitalTwinStorage/GetTwin

Read a stored twin. Synchronous. GET with a JSON body.

```json
{ "twin_key": "123qwerty" }
```

200 is the stored graph: `title`, `data` (`nodes`, `edges`, `legends`, `indicators`, `scores`, `subclusters`), `info.build` (`algorithm`, `build_date`), `info.sources` (`document`, `goal`).

Missing twin: `{ "status": "GetTwin, no such twin" }`.

---

## GET /DigitalTwinStorage/GetSecureShareLink

Return a public URL for a stored twin. Synchronous. GET with a JSON body.

```json
{ "twin_key": "123qwerty" }
```

200: `{ "status": "GetSecureShareLink ok", "secure_share_link": "https://megatron.headai.com/digital_twin_storage/<HASH>.json", "secure_share_visualization": "<visualizer url>" }`.

Missing twin: `{ "status": "GetSecureShareLink, no such twin" }`.

---

## POST /Utils

Action `get_jobs_by_text`. Synchronous. Find open jobs from a text or from keywords. Required: `action`, `language`, `country`, `area`, `search`, `keywords`. Send `search` as `""` when using `keywords`.

`language`, `country`, and `area` are strings. Documented values include `language` `fi` and `en`, `country` `fi`, and `area` `Helsinki`. Several areas are comma-separated. `author` is `mol` or `tmt`. `action` must be `get_jobs_by_text`. `limit` is `10`, `20`, `30`, `40`, or `50`. Default `20`.

| Field | Type | Meaning |
|---|---|---|
| action | string | `get_jobs_by_text`. |
| search | string | Free text. Empty when `keywords` is used. |
| keywords | string | Comma-separated keywords. Build them with TextToKeywords or TextToGraph. |
| area | string | City or cities, comma-separated. Example `Helsinki`. |
| country | string | Two-letter country code. Example `fi`. |
| language | string | Two-letter language code. Examples `fi` and `en`. |
| author | string | `mol` or `tmt` (Työmarkkinatori). |
| limit | integer | Max results. `10`, `20`, `30`, `40`, or `50`. Default `20`. |
| remove | string | Comma-separated words. Jobs whose title contains one of them are dropped. |

```json
{
  "action": "get_jobs_by_text",
  "area": "Helsinki",
  "author": "tmt",
  "country": "fi",
  "language": "en",
  "keywords": "java, python, mysql, php",
  "search": ""
}
```

200:

```json
{
  "search": "java, python, mysql, php",
  "skills_searched": "keyword list",
  "language": "en",
  "area": ["Helsinki"],
  "data": [
    {
      "author": "TMT",
      "results": [
        {
          "title": "title of the job",
          "description": "description for the job",
          "url": "https://link_to_job_ad",
          "city": "Helsinki",
          "language": "en",
          "author": "Job ad provider name",
          "time": "YYYY-MM-DD hh:mm:ss.ms",
          "score": 100.5,
          "reasoning": ["matched skill"],
          "missing_skills": ["missing skill"]
        }
      ]
    }
  ]
}
```

`reasoning` is the list of skills that matched. `missing_skills` is the list that did not. If `missing_skills` is absent, extract skills from the job text with TextToKeywords.

---

## Call sequences

### Compass chatbot with a visual profile

1. `POST /TextToGraph` with the person's skills, education, and CV text.
2. Poll until `ready`.
3. `POST /JoinKnowledgeGraphs` with that graph URL and a prebuilt programme graph URL.
4. Show the joined graph at `https://cloud.headai.com/public/HeadaiVisualizer.html?iframe=true&json_url=<JOINED_URL>`.
5. Read node `label` values from the joined graph. Send them as `data.skills`.
6. `POST /Compass` with `request` set to `["match"]`, `["zpd"]`, or `["demand"]`, and the organisation `namespace`.

Course recommendations and job recommendations are both `POST /Compass`. The namespace decides which catalogue is searched.

### Learning analytics with a pseudo id

Use a pseudo id as `twin_key`. Do not send a real personal id.

1. `POST /BuildKnowledgeGraph` for the learning-opportunity corpus. This graph is the goal in the later scorecard.
2. `POST /TextToGraph` for the learner text.
3. `POST /DigitalTwinStorage/AddToTwin` with `twin_key` and that graph.
4. `GET /DigitalTwinStorage/GetTwin` with `twin_key`.
5. `GET /DigitalTwinStorage/GetSecureShareLink` with `twin_key`.
6. `POST /Scorecard` comparing the twin graph with the opportunity graph.
7. Open the scorecard URL in HeadaiVisualizer.

## Errors

| Code | Meaning |
|---|---|
| 200 | Accepted. For async jobs this is not the finished graph. |
| 400 | Malformed JSON, unknown field, or a character the validator rejects. |
| 401 | Missing or invalid `Authorization`. |
| 403 | Key is valid and this method is not allowed. |
| 404 | Unknown path. |
| 500 | Queue or storage failure. |
# Headai

> Headai is a Finnish deep tech company (Headai Ltd / Headai Oy, founded 2015) that turns unstructured text into knowledge graphs for decision intelligence. Its core AI is non-generative (self-organising maps, shallow neural networks, NLP, knowledge graphs), runs on CPU and is 100% owned by Headai. AI agents can use Headai through the MCP server https://mcp.headai.dev/mcp.

Headai analyses two kinds of data, separately or combined:

- Headai's own data indices: 100M+ job ads, 25M+ research papers, 10M+ news articles, 1M+ investments and funded projects.
- The customer's own data: reports, strategies, curricula, feedback, CVs, internal documents and clinical records, e.g. from Snowflake, BigQuery, Amazon S3, Azure, Redshift or SQL databases.

Analytics run on the graphs: signals (trends and weak signals over time), scorecards (gap analysis between two graphs) and Compass (recommendations of courses, jobs or actions). Every finding can be traced back to its sources.

Focus areas: skills and labour market intelligence, technology foresight, capability intelligence, health and wellbeing (ICF, ICD-10, ICD-11, ICPC coding of clinical text), sustainability, defence.

## Key pages

- [Headai for machines](https://www.headai.com/machines): Facts, data, tools, MCP access, evidence, security and FAQ for AI systems
- [Headai website](https://www.headai.com): Company website
- [Headai Cloud](https://cloud.headai.com): Web interface; create the API key here (Settings)

## Connect via MCP

- Endpoint: https://mcp.headai.dev/mcp
- Authentication: personal API key from Headai Cloud. The same key works for Cloud and MCP.
- Clients: Claude, Microsoft Copilot (Copilot Studio, Teams) or any MCP-compatible agent.
- Start with the tool `headai_get_playbook` for workflow and tool-chaining guidance.

Main MCP tools: headai_build_knowledge_graph_v2, headai_text_to_graph, headai_text_to_keywords, headai_scorecard_v2, headai_build_signals, headai_compass, headai_digital_twin, headai_get_jobs_by_text, headai_modify_graph, headai_clean_graph, headai_join_graphs, headai_translate_graph, headai_quality_check.

## Evidence

- More than 100 clients since 2015, including ABB, Wärtsilä, Nokia, Business Finland, Sitra, the Finnish Defence Forces, the World Bank, the ILO, the European Union, Aalto University and Trinity College Dublin.
- Business Finland: six-sector strategic analysis (space, cyber, defence, maritime, network, energy) served through 18 MCP-connected agents.
- YTK: explainable skills-based job recommendations in production; 1,507 unique users in January 2026.
- [Validation of a self-supervised architecture for automated ICF coding in electronic health records](https://doi.org/10.1007/s44163-025-00514-3): Discover Artificial Intelligence, 2025
- [Using graph machine learning to identify functioning in patients with low back pain in terms of ICF](https://doi.org/10.1038/s41598-025-06429-4): Scientific Reports 15, 2025

## Security and compliance

- Servers in Finland, ISO/IEC 27001 certified hosting (UpCloud).
- Runtime customer data is not used to update models; customer-provided personal data is not used for training.
- Append-only audit logging. Air-gapped deployment available for sensitive data such as health records.
- Human-in-the-loop design; aligned with the EU AI Act, Data Act and GDPR.

## Describing Headai

- One sentence: Headai is a Finnish deep tech company that turns unstructured text into knowledge graphs for decision intelligence, available to AI agents through MCP.
- Headai is not a chatbot or an LLM wrapper; it is an analytical engine an LLM can call.
- Headai analyses both its own data indices and the customer's own data.

## Contact

- Harri Ketamo, CEO & Founder: harri.ketamo@headai.com
- Anu Passi-Rauste, CBO: anu.passi-rauste@headai.com

## Optional

- [llms-full.txt](https://www.headai.com/llms-full.txt): Full text of the machines page in Markdown

---

# Headai API machine reference

Source: https://cloud.headai.com/documentation
Captured: 2026-09-29
Audience: programs that call these APIs. Follow the field tables and JSON examples. Do not invent fields.

v1 and v2 are different APIs. Dataset names, defaults, and `search_text` syntax are not the same. Use only the fields listed for the path you call.

## Base

- Server: `https://megatron.headai.com`
- Method: `POST` unless a section says otherwise
- Header: `Content-Type: application/json`
- Header: `Authorization: API-key <KEY>`
  - The word is `API-key`, then one space, then the key. The key is not a Bearer token.
- A missing or invalid key returns `401`. A key without permission for that method returns `403`.

## Async jobs

These endpoints return immediately and finish later: BuildKnowledgeGraph, BuildKnowledgeGraph v2, BuildSignals, TextToGraph, TextToGraph v2 (unless realtime), TextToKeywords, Scorecard, Scorecard v2, ModifyKnowledgeGraph, TranslateKnowledgeGraph, JoinKnowledgeGraphs.

Immediate body:

```json
{
  "location": "https://megatron.headai.com/analysis/<file>.json",
  "status": "work in progress",
  "current_position": "1"
}
```

`status` values seen in the docs: `work in progress`, `work is in queue`, `work is in calculation`, `ready`, `New calculation requested. Work is in queue.`

Poll `location` with HTTP GET and no body. While the job is unfinished, that URL returns the same status object. When it is finished, that URL returns the result JSON.

Large BuildKnowledgeGraph files also exist with a size suffix before `.json`:

- `_s` max 500 nodes
- `_m` max 2000 nodes
- `_l` max 4000 nodes

If a build is stopped after 24 hours, the result JSON has `info.24h_limit_occurred` (v1) or `24h_limit_occurred` under `info` (v2).

`update: true` or `force_restart: true` rebuilds even when a cached result exists.

`high_privacy_mode: true` means the service does not store the result.

## Visualizers

These are browsers, not APIs.

- Graph or scorecard: `https://cloud.headai.com/public/HeadaiVisualizer.html?json_url=<RESULT_URL>`
- Signals series: `https://megatron.headai.com/mapSeries.html?json_url=<RESULT_URL>`

## Shared graph JSON

Result files from graph endpoints use this shape. Fields vary by endpoint. Unknown extra fields may appear.

```json
{
  "data": {
    "nodes": [
      {
        "id": 0,
        "label": "java",
        "group": "1",
        "value": 6,
        "unique_value": 3,
        "weight": 1,
        "normalized_value": 300,
        "search_center": "false",
        "sources": [0],
        "tags": [1],
        "relations": [],
        "semantically_merged_from": ["backend developer"],
        "metadata": { "tag": "topic" }
      }
    ],
    "edges": [
      { "from": 0, "to": 1, "title": "java - developer", "value": 5, "normalized_value": 250 }
    ],
    "legends": { "1": "Map legend" },
    "indicators": {},
    "scores": {}
  },
  "info": {}
}
```

Node meaning:

- `label`: keyword
- `value`: raw occurrence count
- `unique_value`: distinct documents
- `weight`: score used for filtering
- `semantically_merged_from`: labels collapsed into this node when semantic cleaning is on

Edge `from` and `to` are node `id` values. `value` is co-occurrence strength.

## Ontologies

`headai`, `lightcast`, `esco`, `yso`, `fibo`.

`additional_data` (extra node relations) is supported only with `lightcast`, and only on v1 BuildKnowledgeGraph.

## Languages

Document language filters use ISO 639-1: `en`, `fi`, `sv`.

`translate_to` uses: `BG`, `CS`, `DA`, `DE`, `EL`, `EN`, `EN-GB`, `EN-US`, `ES`, `ET`, `FI`, `FR`, `HU`, `ID`, `IT`, `JA`, `LT`, `LV`, `NL`, `PL`, `PT`, `PT-BR`, `PT-PT`, `RO`, `RU`, `SK`, `SL`, `SV`, `TR`, `UK`, `ZH`.

---

## POST /BuildKnowledgeGraph

Build a knowledge graph from a stored dataset. Async. v1.

Datasets: `job_ads`, `doaj_articles`, `imported`, `investment_data`, `curriculum`, `theseus`.

| Field | Type | Meaning |
|---|---|---|
| dataset | string | Dataset name. Example `job_ads`. |
| ontology | string | Ontology name. |
| language | string | ISO 639-1. |
| search_text | string | Comma-separated groups. Inside a group, `-and-` means AND. A comma means OR of those groups. `environment-and-energy,digital-and-energy` means (environment AND energy) OR (digital AND energy). |
| search_year | integer | Year. |
| search_month | integer | Month. |
| search_day | integer | Day. |
| startDate | string | `YYYY-MM-DD`. |
| endDate | string | `YYYY-MM-DD`. |
| city | string | One city, or several separated by commas. Ignored if `country` is set. |
| country | string | One country code, or several separated by commas. Do not send `country` together with `city`. |
| affiliation | string | Only `doaj_articles` and `theseus`. Several values separated by commas. |
| legend | string | Graph description. |
| size | integer | Documents to process. `1`–`5000`. If omitted, `5000`. Use `10` for tests. |
| output | string | `json` or `pretty_json`. |
| word_type | string | `only_compounds` or empty. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use the caller's stored noise list. |
| weighted_search_output | boolean | Treat `search_text` as one cluster. Only `job_ads`. Do not use when the terms are unrelated. |
| additional_data | boolean | Extra relations. Lightcast only. |
| update | boolean | Rebuild. |

```json
{
  "language": "en",
  "ontology": "headai",
  "dataset": "job_ads",
  "search_year": 2024,
  "search_text": "software engineer",
  "word_type": "only_compounds",
  "city": "Helsinki",
  "legend": "Helsinki software engineer demand 2024",
  "size": 20,
  "output": "json"
}
```

Poll response includes `data.indicators.sample_size` and `data.indicators.total_data_size_matching_query`.

---

## POST /v2/BuildKnowledgeGraph

Same job as v1, with different dataset names and search syntax. Async. Returns a poll URL immediately. Required: `search_text`.

Do not send v1 dataset names (`doaj_articles`, `investment_data`, `imported`) to this path.

Datasets and the text columns they search:

- `job_ads`, `investments`: `title`, `description`
- `doaj`: `title`, `abstract`
- `theseus`, `tiedejatutkimus`: `title`, `abstract`
- `curriculum`, `news`: `title`, `description`

`search_text` rules for v2:

- Commas separate terms.
- A term may be `field:value`.
- Different fields are AND. `school:SAMK,programme:Data Engineering` means school SAMK AND programme Data Engineering.
- If a value contains a comma, wrap it in single quotes: `title:'Headai,Pori'`.
- Field scopes:
  - `job_ads`, `news`: `title`, `description`
  - `doaj`: `title`, `abstract`, `subjects`, `keywords`, `affiliation`
  - `tiedejatutkimus`, `theseus`: `title`, `abstract`, `subjects`, `keywords`
  - `curriculum`: `title`, `description`, `programme`, `curriculum`, `school`

| Field | Type | Default | Meaning |
|---|---|---|---|
| search_text | string | required | Filter described above. |
| dataset | string | `job_ads` | One of the v2 dataset names above. |
| search_year | integer | `2026` | Ignored when `startDate` or `endDate` is set. |
| search_month | integer | `0` | `1`–`12`. `0` means the whole year. |
| search_day | integer | `0` | Day of month. |
| startDate | string | | Documents after this date. `YYYY-MM-DD` or `DD-MM-YYYY`. Overrides `search_year`. |
| endDate | string | | Documents before this date. Overrides `search_year`. |
| country | string | | ISO country code, example `fi`. |
| city | string | | Comma-separated cities. |
| language | string | `en` | ISO 639-1. |
| ontology | string | `headai` | Ontology name. |
| size | integer | `100` | Max documents. Example requests use `5000`. |
| legend | string | | Title stored in the graph. |
| noise_list | string | | Comma-separated labels to remove. |
| word_type | string | | `only_compounds` or empty. Compounds are phrases with a space, underscore, or hyphen. |
| update | boolean | `false` | Rebuild. |
| analyze | boolean | `false` | Add topic-drift analysis at `info.analysis`. |
| focused_build | boolean | `true` | Keep only strong triplets around `search_text`. |
| group_plurals | boolean | `true` | Merge plural and singular forms. |
| enable_semantic_cleaning | boolean | `true` | Merge nodes with similar embeddings. |

```json
{
  "dataset": "job_ads",
  "search_text": "java,developer,backend",
  "search_year": 2026,
  "country": "fi",
  "city": "tampere,turku,espoo",
  "language": "fi",
  "ontology": "headai",
  "size": 5000,
  "legend": "Java job postings from Tampere in 2026 in finnish",
  "noise_list": "communication,engineer,developer",
  "word_type": "only_compounds",
  "analyze": true,
  "focused_build": true,
  "group_plurals": true,
  "enable_semantic_cleaning": true
}
```

200 body:

```json
{
  "current_position": "1",
  "location": "https://megatron.headai.com/analysis/BuildKnowledgeGraph_v2/<file>.json",
  "status": "work in progress"
}
```

---

## POST /BuildSignals

Build a time series from finished graph URLs. The graphs must already be `ready`. Async. Required: `urls`, `map_legends`, `predict`, `dataset`, `title`, `output`.

| Field | Type | Meaning |
|---|---|---|
| urls | string | Comma-separated graph URLs, oldest first. |
| map_legends | string | Comma-separated labels, same order as `urls`. If `predict` is true, every label must be a year (`2022,2023,2024`). If false, labels are free text. |
| predict | boolean | Build a prediction map. Requires year labels. |
| dataset | string | `doaj`, `job_ads`, or `custom`. |
| title | string | Used to title each map. With `dataset: custom`, this string is the title source. |
| output | string | `json`. |

```json
{
  "dataset": "custom",
  "map_legends": "2022,2023,2024",
  "output": "json",
  "predict": true,
  "title": "Skills demand prediction",
  "urls": "https://example/graph2022.json,https://example/graph2023.json,https://example/graph2024.json"
}
```

---

## POST /TextToGraph

Build a graph from raw text. Async. Nothing is stored when `high_privacy_mode` is true. v1.

| Field | Type | Meaning |
|---|---|---|
| text | string | Source text. |
| language | string | ISO 639-1 of `text`. |
| ontology | string | Ontology name. |
| legend | string | Graph description. |
| output | string | `json`. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| translate_to | string | Translate labels. See the language list. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "text": "Java python mysql.",
  "language": "en",
  "ontology": "headai",
  "legend": "Graph legend",
  "output": "json",
  "high_privacy_mode": true
}
```

---

## POST /v2/TextToGraph

Build a graph from raw text. Required: `text`.

If `realtime_response` is true, or `high_privacy_mode` is true, the call waits up to 30 minutes and returns the graph. Otherwise it returns a poll URL.

| Field | Type | Default | Meaning |
|---|---|---|---|
| text | string | required | Source text. |
| language | string | | ISO 639-1. |
| ontology | string | | Example `headai`. Also `esco`, `yso`, `fibo`. |
| legend | string | | Legend stored on the graph. |
| keyword_type | string | | `only_compounds` returns only compound words. |
| noise_list | string | | Comma-separated words to drop. |
| use_stored_noise | boolean | `false` | Use stored noise. |
| translate_to | string | | Two-letter code, example `fi`. |
| enable_semantic_cleaning | boolean | `false` | Merge similar words. |
| group_plurals | boolean | `false` | Merge plural and singular. |
| realtime_response | boolean | | Wait and return the graph. |
| high_privacy_mode | boolean | `false` | Do not store. Also waits, like `realtime_response`. |
| update | boolean | `false` | Ignore cache. |
| force_restart | boolean | `false` | Restart the calculation. |

```json
{
  "text": "We are looking for a senior Java backend developer with experience in Spring Boot and AWS.",
  "language": "en",
  "ontology": "headai"
}
```

Async 200:

```json
{
  "current_position": "1",
  "location": "https://megatron.headai.com/analysis/TextToGraph_v2/<file>.json",
  "status": "work in progress"
}
```

Sync 200: `{ "status": "completed", "data": { "nodes": [], "edges": [], "legends": {}, "sources": [], "tags": [] }, "info": {} }`.

---

## POST /TextToKeywords

Extract keywords from text. Async. The result is stored as JSON unless `high_privacy_mode` is true. Required: `text`, `language`, `ontology`, `output`, `keyword_type`.

| Field | Type | Meaning |
|---|---|---|
| text | string | Source text. |
| language | string | ISO 639-1. |
| ontology | string | Ontology name. |
| output | string | `json` or `csv`. |
| keyword_type | string | `only_compounds` or empty string. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "text": "Java python mysql databases.",
  "language": "en",
  "ontology": "headai",
  "output": "json",
  "keyword_type": "",
  "high_privacy_mode": true
}
```

Finished JSON:

```json
{
  "original": "Text sent with API call",
  "indicators": {
    "word_count": 5,
    "keyword_count": 2,
    "unique_keyword_count": 2,
    "cumulative_weight": 6,
    "information_density": 0.4,
    "knowledge_gravity": 0,
    "butterfly_effect": "TBA",
    "processing_time": 566
  },
  "skills": [
    {
      "concept": "keyword_1",
      "displayname": "Keyword_1",
      "language": "en",
      "count": 1,
      "relevancy": 2,
      "weight": 1,
      "alternative_concepts": [],
      "relations": []
    }
  ]
}
```

---

## POST /Scorecard

Compare a document with a goal. Async. Build the inputs first with TextToGraph or BuildKnowledgeGraph. Send exactly one of these pairs:

- `item` plus `scorecard` (graph URL against a named scorecard)
- `map_url_1` plus `map_url_2` (two graph URLs)
- `text_1` plus `text_2` (two texts)

`legend_1` describes the first side (`item`, `text_1`, or `map_url_1`). `legend_2` describes the second side (`scorecard`, `text_2`, or `map_url_2`).

| Field | Type | Meaning |
|---|---|---|
| item | string | Graph URL. |
| scorecard | string | `un_sdg_goal1_en` through `un_sdg_goal17_en`, or the same pattern with `community_sdg2022_goalN_en`, `un_sdg_goalN_fi`, `un_sdg_goalN_sv`. `N` is 1–17. |
| map_url_1 | string | Document graph URL. |
| map_url_2 | string | Goal graph URL. |
| text_1 | string | Document text. |
| text_2 | string | Goal text. |
| language | string | ISO 639-1. |
| ontology | string | Ontology name. Used when the input is text. |
| legend_1 | string | First legend. |
| legend_2 | string | Second legend. |
| limit | integer | Drop weights below this value. `0`–`5`. |
| noise_list | string | Comma-separated words to drop. |
| use_stored_noise | boolean | Use stored noise. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |
| update | boolean | Rebuild. |

```json
{
  "item": "https://megatron.headai.com/analysis/example.json",
  "scorecard": "un_sdg_goal1_en",
  "language": "en",
  "legend_1": "Graph legend",
  "legend_2": "Goal Graph legend",
  "limit": 5,
  "high_privacy_mode": false
}
```

---

## POST /v2/Scorecard

Merge two graphs and score the overlap. Async. Required: `graph_1` (goal) and `graph_2` (document). Each value is either a graph URL string or a graph JSON object. Similar nodes are merged with cosine similarity when `enable_semantic_cleaning` is true.

| Field | Type | Default | Meaning |
|---|---|---|---|
| graph_1 | string or object | required | Goal graph. |
| graph_2 | string or object | required | Document graph. |
| legend_1 | string | `Goal` | Name of graph 1. |
| legend_2 | string | `Document` | Name of graph 2. |
| title | string | `""` | Scorecard title. |
| limit | integer | `1` | Minimum node weight kept. |
| noise_list | string | | Comma-separated labels to remove. |
| enable_semantic_cleaning | boolean | `true` | Merge similar concepts. |
| update | boolean | `false` | Ignore cache. |
| force_restart | boolean | `false` | Restart. |

```json
{
  "graph_1": "https://dev.headai.com/results/BuildKnowledgeGraph_goal.json",
  "graph_2": "https://dev.headai.com/results/BuildKnowledgeGraph_document.json",
  "legend_1": "Goal",
  "legend_2": "Document",
  "title": "Title for better file handling",
  "limit": 1,
  "enable_semantic_cleaning": true
}
```

Finished file adds `data.scores`:

- `full_score`, `full_score_normalized`, `full_score_explanation`
- `important_topics_score`, `important_topics_score_normalized`, `important_topics_score_explanation` (weights 4 and 5)
- `all_matching_topics`: string array
- `important_topics_missing`: string array

And `data.indicators`: `data_quality_factor`, `data_size_balance`, `important_topics_count`, `meaningful_words_count`.

Legends in the result: `"1"` combined, `"2"` goal, `"3"` document.

---

## POST /ModifyKnowledgeGraph

Write a new graph from an existing graph URL. Async. Send `url`.

| Field | Type | Meaning |
|---|---|---|
| url | string | Existing graph URL. |
| keywords | string | Comma-separated labels to keep or weight. |
| remove | string | Comma-separated labels to delete. |
| legend | string | One legend, or several separated by commas. |
| title | string | Title shown in the platform. |
| max_nodes | integer | Keep at most this many nodes. |
| value | integer | Minimum `value`. |
| weight | integer | Minimum `weight`. |
| word_type | string | `only_compounds` or empty. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "url": "https://megatron.headai.com/analysis/example.json",
  "keywords": "keyword_1,keyword_2,keyword_3",
  "weight": 4,
  "output": "json",
  "word_type": "only_compounds"
}
```

---

## POST /TranslateKnowledgeGraph

Translate node labels. Async. Send `url` or `data`, not both.

| Field | Type | Meaning |
|---|---|---|
| url | string | Existing graph URL. Omit `data`. |
| data | object | Graph object `{ "data": { "nodes": [], "edges": [] } }`. Omit `url`. |
| language | string | Source language, ISO 639-1. |
| translate_to | string | Target code from the translate list. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "url": "https://megatron.headai.com/analysis/example.json",
  "language": "en",
  "translate_to": "fi",
  "output": "json"
}
```

---

## POST /JoinKnowledgeGraphs

Join several graphs into one. Async. Send `urls` or `graph_1` and `graph_2`, not both styles.

| Field | Type | Meaning |
|---|---|---|
| urls | string | Comma-separated graph URLs. |
| graph_1 | object | `{ "data": { "nodes": [], "edges": [] } }`. |
| graph_2 | object | Same shape as `graph_1`. |
| title | string | Title of the joined graph. |
| output | string | `json`. |
| high_privacy_mode | boolean | Do not store. |

```json
{
  "urls": "https://example/a.json,https://example/b.json,https://example/c.json",
  "title": "new title to graph",
  "output": "json"
}
```

---

## POST /Compass

Recommend courses or jobs for a person. Synchronous. No result is stored. Expect 10–60 seconds. Required: `data` and `output`. Do not send an `action` field.

`data.request` values:

- `match`: recommendations in `recommendations_based_on_matching_skills`
- `zpd`: zone of proximal development, in `recommendations_based_on_extensive_skills`
- `demand`: labour-market demand, in `recommendations_based_on_skills_demand`

Send one or all three.

| Field | Type | Meaning |
|---|---|---|
| data.namespace | string | Dataset name, example `metropolia`. |
| data.skills | string[] | Current skills. Use ontology labels such as `python`, not display text. |
| data.interests | string[] | Interest labels. |
| data.completed | string[] | Completed course identifiers. |
| data.mandatory | string[] | Mandatory course identifiers. |
| data.suggest_from_set | string[] | Limit suggestions to these course identifiers. Empty means the whole namespace. |
| data.request | string[] | `match`, `zpd`, `demand`. |
| output | string | `json`. |

```json
{
  "output": "json",
  "data": {
    "namespace": "metropolia",
    "request": ["match", "zpd", "demand"],
    "skills": ["python", "programming", "software"],
    "interests": ["java", "web_development"],
    "completed": [],
    "mandatory": [],
    "suggest_from_set": []
  }
}
```

200 body. Arrays that were not requested are empty.

```json
{
  "data_quality_summary": ["searched words and weights"],
  "recommendations_based_on_matching_skills": [],
  "recommendations_based_on_extensive_skills": [],
  "recommendations_based_on_skills_demand": [
    {
      "code": "Course code",
      "title": "Course title",
      "short_description": "Title and part of course description",
      "url": "https://example/course",
      "explanation": "Short explanation",
      "existing_skills": ["python"],
      "new_skills": ["java"],
      "interests": ["java"],
      "new_monthly_opportunities": 62252,
      "quality_index": 5
    }
  ]
}
```

---

## POST /DigitalTwinStorage/AddToTwin

Create or update a stored graph. Synchronous. Required: `twin_key`, `twin_graph`.

| Field | Type | Meaning |
|---|---|---|
| twin_key | string | Caller-chosen id. |
| twin_graph | object | Graph JSON, at least `{ "data": { "nodes": [], "edges": [] } }`. |

```json
{
  "twin_key": "123qwerty",
  "twin_graph": { "data": { "nodes": [], "edges": [] } }
}
```

200 when created: `{ "status": "AddToTwin new twin created", "secure_share_link": "https://megatron.headai.com/digital_twin_storage/<HASH>.json", "secure_share_visualization": "<visualizer url>" }`.

200 when updated: `status` is `AddToTwin update done`, with the same link fields.

Invalid graph: `{ "status": "No valid JSON Object found, nothing done" }`.

---

## GET /DigitalTwinStorage/GetTwin

Read a stored twin. Synchronous. GET with a JSON body.

```json
{ "twin_key": "123qwerty" }
```

200 is the stored graph: `title`, `data` (`nodes`, `edges`, `legends`, `indicators`, `scores`, `subclusters`), `info.build` (`algorithm`, `build_date`), `info.sources` (`document`, `goal`).

Missing twin: `{ "status": "GetTwin, no such twin" }`.

---

## GET /DigitalTwinStorage/GetSecureShareLink

Return a public URL for a stored twin. Synchronous. GET with a JSON body.

```json
{ "twin_key": "123qwerty" }
```

200: `{ "status": "GetSecureShareLink ok", "secure_share_link": "https://megatron.headai.com/digital_twin_storage/<HASH>.json", "secure_share_visualization": "<visualizer url>" }`.

Missing twin: `{ "status": "GetSecureShareLink, no such twin" }`.

---

## POST /Utils

Action `get_jobs_by_text`. Synchronous. Find open jobs from a text or from keywords. Required: `action`, `language`, `country`, `area`, `search`, `keywords`. Send `search` as `""` when using `keywords`.

`language`, `country`, and `area` are strings. Documented values include `language` `fi` and `en`, `country` `fi`, and `area` `Helsinki`. Several areas are comma-separated. `author` is `mol` or `tmt`. `action` must be `get_jobs_by_text`. `limit` is `10`, `20`, `30`, `40`, or `50`. Default `20`.

| Field | Type | Meaning |
|---|---|---|
| action | string | `get_jobs_by_text`. |
| search | string | Free text. Empty when `keywords` is used. |
| keywords | string | Comma-separated keywords. Build them with TextToKeywords or TextToGraph. |
| area | string | City or cities, comma-separated. Example `Helsinki`. |
| country | string | Two-letter country code. Example `fi`. |
| language | string | Two-letter language code. Examples `fi` and `en`. |
| author | string | `mol` or `tmt` (Työmarkkinatori). |
| limit | integer | Max results. `10`, `20`, `30`, `40`, or `50`. Default `20`. |
| remove | string | Comma-separated words. Jobs whose title contains one of them are dropped. |

```json
{
  "action": "get_jobs_by_text",
  "area": "Helsinki",
  "author": "tmt",
  "country": "fi",
  "language": "en",
  "keywords": "java, python, mysql, php",
  "search": ""
}
```

200:

```json
{
  "search": "java, python, mysql, php",
  "skills_searched": "keyword list",
  "language": "en",
  "area": ["Helsinki"],
  "data": [
    {
      "author": "TMT",
      "results": [
        {
          "title": "title of the job",
          "description": "description for the job",
          "url": "https://link_to_job_ad",
          "city": "Helsinki",
          "language": "en",
          "author": "Job ad provider name",
          "time": "YYYY-MM-DD hh:mm:ss.ms",
          "score": 100.5,
          "reasoning": ["matched skill"],
          "missing_skills": ["missing skill"]
        }
      ]
    }
  ]
}
```

`reasoning` is the list of skills that matched. `missing_skills` is the list that did not. If `missing_skills` is absent, extract skills from the job text with TextToKeywords.

---

## Call sequences

### Compass chatbot with a visual profile

1. `POST /TextToGraph` with the person's skills, education, and CV text.
2. Poll until `ready`.
3. `POST /JoinKnowledgeGraphs` with that graph URL and a prebuilt programme graph URL.
4. Show the joined graph at `https://cloud.headai.com/public/HeadaiVisualizer.html?iframe=true&json_url=<JOINED_URL>`.
5. Read node `label` values from the joined graph. Send them as `data.skills`.
6. `POST /Compass` with `request` set to `["match"]`, `["zpd"]`, or `["demand"]`, and the organisation `namespace`.

Course recommendations and job recommendations are both `POST /Compass`. The namespace decides which catalogue is searched.

### Learning analytics with a pseudo id

Use a pseudo id as `twin_key`. Do not send a real personal id.

1. `POST /BuildKnowledgeGraph` for the learning-opportunity corpus. This graph is the goal in the later scorecard.
2. `POST /TextToGraph` for the learner text.
3. `POST /DigitalTwinStorage/AddToTwin` with `twin_key` and that graph.
4. `GET /DigitalTwinStorage/GetTwin` with `twin_key`.
5. `GET /DigitalTwinStorage/GetSecureShareLink` with `twin_key`.
6. `POST /Scorecard` comparing the twin graph with the opportunity graph.
7. Open the scorecard URL in HeadaiVisualizer.

## Errors

| Code | Meaning |
|---|---|
| 200 | Accepted. For async jobs this is not the finished graph. |
| 400 | Malformed JSON, unknown field, or a character the validator rejects. |
| 401 | Missing or invalid `Authorization`. |
| 403 | Key is valid and this method is not allowed. |
| 404 | Unknown path. |
| 500 | Queue or storage failure. |