Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 4 additions & 7 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -129,13 +129,10 @@ so the version a reader picks in the navbar is always one they can install.
A snapshot is changed only where it was wrong for its own release, a patch release that changes
behaviour included. A new feature goes into `docs/`, never back into a snapshot.

**Where this stands.** No snapshot exists yet, so the published site describes `main`, and a
reader on 0.2.0 can meet calls their version lacks; the quick start note warns them. The steps
above are all that is left to run, and `editCurrentVersion` is already set so a snapshot's "Edit
this page" will point at `docs/`. The first cut waits on the first tagged release; the SDK has
tagged only `v0.2.0`, and bumped `main` past it without releasing, so no number in between gets a
snapshot. Nothing checks these rules automatically yet. The doc tests being added under
`doctests/` are meant to, by building against pinned SDK and platform commits.
**Where this stands.** The `1.0` snapshot was cut when Python, Rust and Java all released 1.0.0.
It is served at `/` and `docs/` is served at `/next/`. The next cut waits on the next tagged `vX.Y.0`.
Nothing checks these rules automatically yet. The doc tests being added under `doctests/` are meant
to, by building against pinned SDK and platform commits.

## Every example has to be runnable

Expand Down
16 changes: 8 additions & 8 deletions docs/quickstart.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -11,15 +11,15 @@ From zero to a stored datapoint in about five minutes.

## 1. Install

:::note Which version these pages describe
These pages follow the SDKs as they are being developed, which can be ahead of the latest
release. The Python and Rust SDKs are pre-1.0, so a release can still rename or remove a call.
If an example calls something your installed SDK does not have, it may be newer than your release.
:::note Which version this page describes
This is the Next (unreleased) version of the documentation. It follows `main` of the SDKs, so it can
use calls that the latest release does not have. Pick the release you have installed in the version
menu in the navbar.

| SDK | Latest release |
| --- | --- |
| Python | 0.2.0, on PyPI |
| Rust | 0.2.0, on crates.io |
| Python | 1.0.0, on PyPI |
| Rust | 1.0.0, on crates.io |
| Java | 1.0.0, on Maven Central |
:::

Expand Down Expand Up @@ -58,7 +58,7 @@ pip install intellistream-datahub-sdk
```toml
# Cargo.toml
[dependencies]
intellistream-datahub-sdk = "0.2"
intellistream-datahub-sdk = "1.0"
```

</TabItem>
Expand All @@ -75,7 +75,7 @@ import intellistream_datahub_sdk as dh

```toml
# Cargo.toml — a crate-wide rename
dh = { package = "intellistream-datahub-sdk", version = "0.2" }
dh = { package = "intellistream-datahub-sdk", version = "1.0" }
```
:::

Expand Down
4 changes: 2 additions & 2 deletions docs/tutorial.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,7 @@ portably:
```toml
# Cargo.toml
[dependencies]
intellistream-datahub-sdk = "0.2"
intellistream-datahub-sdk = "1.0"
tokio = { version = "1", features = ["full"] }
chrono = "0.4"
sysinfo = "0.37"
Expand All @@ -113,7 +113,7 @@ The code below lives in `src/main.rs`, with `#[tokio::main]` driving the async c
```toml
# Cargo.toml
[dependencies]
intellistream-datahub-sdk = { version = "0.2", features = ["blocking"] }
intellistream-datahub-sdk = { version = "1.0", features = ["blocking"] }
chrono = "0.4"
sysinfo = "0.37"
hostname = "0.4"
Expand Down
18 changes: 7 additions & 11 deletions docusaurus.config.js
Original file line number Diff line number Diff line change
Expand Up @@ -28,11 +28,11 @@ const config = {
docs: {
sidebarPath: './sidebars.js',
routeBasePath: '/', // docs at site root, GitBook-style
// No `versions` block and no snapshot yet: `docs/` is the only version
// and is served at '/'. AGENTS.md has the three lines to add when the
// first release is tagged. Leave `lastVersion` out then too — it
// defaults to the newest name in versions.json, so cutting a snapshot
// moves readers onto it with nothing here to keep in step.
// `docs/` is the development version, served at /next/. The newest
// snapshot in versions.json is the release and is served at '/'.
// Leave `lastVersion` out: it defaults to the newest name in
// versions.json, so cutting a snapshot moves readers onto it.
versions: { current: { label: 'Next (unreleased)' } },

// Adds "Edit this page" to every doc. Docusaurus appends the file's
// path relative to this site directory. Note the branch here is
Expand All @@ -41,7 +41,7 @@ const config = {
// Send every "Edit this page" to docs/, including from a snapshot once
// one exists: a snapshot takes corrections for its own release, but a
// new feature belongs in docs/, and the stock link would invite the
// wrong edit. No effect until the first snapshot.
// wrong edit.
editCurrentVersion: true,
},
blog: false, // SDK docs site — no blog
Expand Down Expand Up @@ -88,11 +88,7 @@ const config = {
// which pushState's the URL and then renders this site's own 404 —
// the href looks right in the HTML but the click never leaves the SPA.
{ type: 'html', position: 'right', value: '<a class="navbar__item navbar__link" href="/data-platform-documentation/">Platform documentation</a>' },
// The site's own version, not an SDK version. Swap it for
// { type: 'docsVersionDropdown', position: 'right' } once a snapshot
// exists — with none, the dropdown renders as a lone link labelled
// after the development version, which reads as a release.
{ type: 'html', position: 'right', value: '<span class="badge badge--secondary navbar__version-badge">v1.0</span>' },
{ type: 'docsVersionDropdown', position: 'right' },
// This site's own repo, so "GitHub" is unambiguous. The old link went
// to the SDK code on Gitea; if a code link is wanted too it needs its
// own item, since the SDK spans three language repos.
Expand Down
9 changes: 9 additions & 0 deletions versioned_docs/version-1.0/advanced/_category_.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
{
"label": "Advanced scenarios (AI)",
"position": 7,
"link": {
"type": "generated-index",
"title": "Advanced scenarios (AI)",
"description": "Longer, end-to-end builds that go past CRUD into prediction and classification. No data-science background needed — start with 'Machine learning, gently' for the ideas in plain language, then 'Generate sample data' to get a sandbox you can run everything against. The SDK gets data in and out; the learning step in between is shown in Python. Each page carries an effort estimate."
}
}
174 changes: 174 additions & 0 deletions versioned_docs/version-1.0/advanced/asset-health-score.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,174 @@
---
sidebar_position: 4
title: Asset health scoring
---
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';

# Asset health scoring

:::info At a glance
**Effort:** ~20–30 minutes · **You'll build:** a composite 0–100 health index from
several signals, with a healthy/watch/critical classification · **Stack:** the SDK plus
a little arithmetic, no model training.
:::

A machine rarely fails on one signal. Vibration is creeping up, the bearing runs a
little hot, oil pressure sags at load, each is fine alone, but together they tell a
story. A **health score** rolls those signals into one comparable number so an operator
can rank a whole fleet at a glance and a dashboard can show green/amber/red. It's the
lightweight cousin of [predictive maintenance](/advanced/predictive-maintenance): no
model to train, just a transparent, tunable index.

:::info Needs a sandbox
This reads several pump signals. [Generate a sandbox](/advanced/generate-sample-data)
first, section A ingests `pump_07_bearing_temp_c`, `pump_07_oil_pressure_kpa` and a
stand-in `pump_07_vibration_anomaly` (or run [predictive maintenance](/advanced/predictive-maintenance)
to produce the real one).
:::

:::tip New to machine learning?
No background needed. Skim the [gentle primer](/advanced/machine-learning-gently) for the
ideas in plain language, model, feature, training, and the algorithm itself.
:::

## 1. Pull the latest value of each signal

Take the most recent reading of each contributing series for the asset. The seed writes
these hourly, so ask for the latest datapoint rather than a short window that may be empty.

<Tabs groupId="lang">
<TabItem value="python" label="Python">

```python
import intellistream_datahub_sdk, pandas as pd

client = intellistream_datahub_sdk.DataHubClient.from_env()

def latest(external_id):
pts = client.timeseries.retrieve_latest_datapoints([external_id])[0].get_datapoints()
return float(pts[-1].value)

signals = {
"vibration": latest("pump_07_vibration_anomaly"), # 0..~1, from the anomaly model
"bearing_temp": latest("pump_07_bearing_temp_c"),
"oil_pressure": latest("pump_07_oil_pressure_kpa"),
}
```

</TabItem>
<TabItem value="java" label="Java">

```java
// Java has no latest-datapoint call: read the last day and take the newest point.
// The scoring arithmetic below is shown in Python.
var filter = new RetrieveFilter();
filter.setExternalId("pump_07_bearing_temp_c");
filter.setStart(ZonedDateTime.now().minusHours(24));
filter.setEnd(ZonedDateTime.now());
filter.setLimit(100);

var request = new DataRetriever<RetrieveFilter>();
request.setItems(List.of(filter));
var pts = client.timeseries().retrieve(request).getItems().get(0).getDatapoints();
double bearingTemp = Double.parseDouble(pts.get(pts.size() - 1).getValue());
```

</TabItem>
<TabItem value="rust" label="Rust">

```rust
// The latest datapoint, per signal; the scoring arithmetic below is shown in Python.
use intellistream_datahub_sdk::generic::{DataWrapper, IdAndExtId};

let page = api.time_series
.retrieve_latest_datapoint(&DataWrapper::from(vec![
IdAndExtId::from_external_id("pump_07_bearing_temp_c")])).await?;
let latest = &page.get_items()[0];
let bearing_temp = latest.datapoints.last().and_then(|p| p.value).unwrap_or(f64::NAN);
```

</TabItem>
</Tabs>

## 2. Normalise each signal to a "badness" in 0–1

Every signal lives on its own scale, so map each to a common 0 (fine) → 1 (alarm) range
against its healthy and limit values. A reading at or below healthy scores 0; at or
above the limit scores 1; in between, it ramps linearly.

```python
def badness(value, healthy, limit):
if limit == healthy:
return 0.0
return max(0.0, min(1.0, (value - healthy) / (limit - healthy)))

LIMITS = { # (healthy, limit) per signal
"vibration": (0.05, 0.30),
"bearing_temp": (60.0, 90.0),
"oil_pressure": (350.0, 250.0), # inverted: lower is worse
}
parts = {k: badness(v, *LIMITS[k]) for k, v in signals.items()}
```

## 3. Weight, combine, and classify

Weight the signals by how much each matters for this asset class, combine into a 0–100
score (100 = perfect health), and bucket it.

```python
WEIGHTS = {"vibration": 0.5, "bearing_temp": 0.3, "oil_pressure": 0.2}

badness_total = sum(parts[k] * WEIGHTS[k] for k in parts)
score = round(100 * (1 - badness_total), 1)

band = "healthy" if score >= 80 else "watch" if score >= 60 else "critical"
```

## 4. Publish the score and flag the bad ones

Write the score back as its own series, now you can chart, rank and subscribe to asset
health like any other signal, and raise an event when an asset drops to `critical`.

```python
client.timeseries.create([intellistream_datahub_sdk.TimeSeries(
external_id="pump_07_health_score", name="Pump 07 health score", unit="score", value_type="float")])
client.timeseries.insert_from_lists(
timestamps=[pd.Timestamp.now(tz="UTC")], values=[score], ts="pump_07_health_score")

if band == "critical":
client.events.create([intellistream_datahub_sdk.Event(
external_id=f"health_critical_pump_07_{int(pd.Timestamp.now().timestamp())}",
type="health_critical", status="open",
event_time=pd.Timestamp.now(tz="UTC"),
metadata={"asset": "pump_07", "score": str(score),
"worst_signal": max(parts, key=parts.get)})])

# read the score back
stored = client.timeseries.retrieve_latest_datapoints(["pump_07_health_score"])[0].get_datapoints()
print(f"health {stored[-1].value:.1f} ({band})")
assert abs(stored[-1].value - score) < 0.01, "the score was not written"
```

Run it across the fleet on a schedule and you have a single ranked health view,
`worst_signal` in the metadata tells maintenance *why* each asset is red.

## Where to take it further

- **Feed it the model.** Swap the hand-set vibration limits for the
[anomaly score](/advanced/predictive-maintenance) as a direct input.
- **Trend the score.** A falling health score over days is itself a predictor, forecast
it like any [other series](/advanced/demand-forecasting).
- **Roll up the graph.** Average child scores up a [site graph](/guides/model-assets-graph)
for a line- or plant-level health number.

## Further reading

- **Normalising signals to a common scale**: [Feature scaling](https://en.wikipedia.org/wiki/Feature_scaling)
- **The ideas in plain language**: [Machine learning, gently](/advanced/machine-learning-gently)

## See also

- [Query & aggregate](/guides/query-and-aggregate): pulling the contributing signals.
- [Predictive maintenance](/advanced/predictive-maintenance): a learned input to the score.
- [Turn readings into events](/guides/detect-events): flagging critical assets.
Loading
Loading