docs: refresh README and 1.6.0 release notes

README:
- The InfluxDB/Prometheus metric backends described in the overview do not
  exist; describe the built-in metric store and the external cc-metric-store
  routing instead.
- The demo instructions still told the reader to add cluster entries to
  config.json. Clusters come from the cluster.json in each job archive
  directory.
- Dead link to docs/JWT-Handling.md (there is no docs/ directory), configs/
  described as holding environment-variable documentation, missing tools, OIDC
  missing from the list of authentication methods, and web/ indented as a child
  of tools/ in the project structure.
- Added a Configuration section: the top-level sections and which are required,
  the "-file" indirection, and how unknown keys are handled per section.

ReleaseNotes: the 1.6.0 entry covered only the session backend and none of the
breaking changes. Added the .env removal, the policy-based resampling config
(configs carrying trigger/resolutions/minimum-points now abort), the switch to
average downsampling, the OIDC role-mapping requirement, LDAP role-filters, the
auth schema now actually validating, and the new features and bug fixes since
v1.5.4. Fixed the clusterockpit.org typo and marked the 1.5.2 known-issues list
as belonging to that release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-27 16:02:50 +02:00
co-authored by Claude Opus 5
parent 6e49b0a332
commit a1e8812547
2 changed files with 187 additions and 20 deletions
+62 -13
View File
@@ -30,10 +30,12 @@ based on a GraphQL API. The web frontend is also served by the backend using
[Bootstrap Icons](https://icons.getbootstrap.com/). [Bootstrap Icons](https://icons.getbootstrap.com/).
The backend uses [SQLite 3](https://sqlite.org/) as the relational SQL database. The backend uses [SQLite 3](https://sqlite.org/) as the relational SQL database.
While there are metric data backends for the InfluxDB and Prometheus time series Time-series metric data is kept separately: cc-backend ships with a built-in
databases, the only tested and supported setup is to use cc-metric-store as the in-memory metric store (configured via the `metric-store` section) that persists
metric data backend. Documentation on how to integrate ClusterCockpit with other its buffers with a write-ahead log and periodic binary checkpoints. Alternatively
time series databases will be added in the future. (or in addition, routed per cluster) it can query one or more external
[cc-metric-store](https://github.com/ClusterCockpit/cc-metric-store) instances via
the `metric-store-external` section. At least one of the two must be configured.
For real-time integration with HPC systems, the backend can subscribe to For real-time integration with HPC systems, the backend can subscribe to
[NATS](https://nats.io/) subjects to receive job start/stop events and node [NATS](https://nats.io/) subjects to receive job start/stop events and node
@@ -42,8 +44,9 @@ state updates, providing an alternative to REST API polling.
Completed batch jobs are stored in a file-based job archive according to Completed batch jobs are stored in a file-based job archive according to
[this specification](https://github.com/ClusterCockpit/cc-specifications/tree/main/job-archive). [this specification](https://github.com/ClusterCockpit/cc-specifications/tree/main/job-archive).
The backend supports authentication via local accounts, an external LDAP The backend supports authentication via local accounts, an external LDAP
directory, and JWT tokens. Authorization for APIs is implemented with directory, an OpenID Connect provider, and JWT tokens. Authorization for APIs is
[JWT](https://jwt.io/) tokens created with public/private key encryption. implemented with [JWT](https://jwt.io/) tokens created with public/private key
encryption.
You find a detailed documentation on the [ClusterCockpit You find a detailed documentation on the [ClusterCockpit
Webpage](https://clustercockpit.org). Webpage](https://clustercockpit.org).
@@ -77,13 +80,16 @@ Execute the following steps:
```shell ```shell
./cc-backend -init ./cc-backend -init
vim config.json (Add a second cluster entry and name the clusters alex and fritz)
wget https://hpc-mover.rrze.uni-erlangen.de/HPC-Data/0x7b58aefb/eig7ahyo6fo2bais0ephuf2aitohv1ai/job-archive-demo.tar wget https://hpc-mover.rrze.uni-erlangen.de/HPC-Data/0x7b58aefb/eig7ahyo6fo2bais0ephuf2aitohv1ai/job-archive-demo.tar
tar xf job-archive-demo.tar tar xf job-archive-demo.tar
./cc-backend -init-db -add-user demo:admin:demo -loglevel info ./cc-backend -init-db -add-user demo:admin:demo -loglevel info
./cc-backend -server -dev -loglevel info ./cc-backend -server -dev -loglevel info
``` ```
The clusters are read from the job archive (one directory with a `cluster.json`
per cluster), so no cluster list has to be maintained in `config.json`. The demo
archive contains the clusters `alex` and `fritz`.
You can access the web interface at [http://localhost:8080](http://localhost:8080). You can access the web interface at [http://localhost:8080](http://localhost:8080).
Credentials for login are `demo:demo`. Credentials for login are `demo:demo`.
Please note that some views do not work without a metric backend (e.g., the Please note that some views do not work without a metric backend (e.g., the
@@ -130,6 +136,8 @@ cd ./cc-backend/
make make
cp configs/config.json . cp configs/config.json .
# config.json references the UI defaults via "ui-file", so copy that too:
cp configs/uiConfig.json .
# EDIT config.json BEFORE YOU DEPLOY: change the secrets under "auth.jwts" # EDIT config.json BEFORE YOU DEPLOY: change the secrets under "auth.jwts"
# ("public-key"/"private-key"). Each secret can also be supplied via an # ("public-key"/"private-key"). Each secret can also be supplied via an
# environment variable (e.g. JWT_PUBLIC_KEY), which takes precedence over the # environment variable (e.g. JWT_PUBLIC_KEY), which takes precedence over the
@@ -172,6 +180,42 @@ For production deployments, serve cc-backend over HTTPS so the session cookie is
marked `Secure`. If you terminate TLS at a reverse proxy, prefer letting marked `Secure`. If you terminate TLS at a reverse proxy, prefer letting
cc-backend serve HTTPS directly for now so the flag is applied. cc-backend serve HTTPS directly for now so the flag is applied.
## Configuration
cc-backend reads a single JSON configuration file (`./config.json` by default,
override with `-config`). Every top-level key is one configuration section:
| Section | Required | Description |
| ----------------------- | -------- | ----------------------------------------------------------------------------------------------------------------- |
| `main` | yes | Server address, TLS, database path and tuning, resampling defaults, footer links, job taggers, retention of node states, NATS API subjects. |
| `auth` | no | Local, LDAP, OIDC and JWT authentication including all secrets. |
| `nats` | no | Connection to the NATS broker (address, credentials). Required for the NATS API and for NATS metric ingestion. |
| `archive` | no | Job archive backend (`file`, `s3` or `sqlite`), plus archive retention and compression. Defaults to `./var/job-archive`. |
| `metric-store` | no* | Built-in in-memory metric store: retention, memory cap, checkpointing, cleanup and NATS subscriptions. |
| `metric-store-external` | no* | List of external cc-metric-store endpoints with a `scope` (cluster name or `*`) used to route queries. |
| `cron` | no | Intervals of the background workers (commit-job, duration, footprint). |
| `ui` | no | Defaults for the web frontend shown to new users (job list, job view, metric selections, plot options). |
\* At least one of `metric-store` and `metric-store-external` must be present,
otherwise the server refuses to start.
Any section can be moved into its own file by appending `-file` to the key and
giving a path instead of an object, e.g. `"ui-file": "uiConfig.json"` loads the
`ui` section from `uiConfig.json`. Unknown top-level keys are ignored with a
warning. Unknown keys *inside* a section are always rejected by the parser: for
`main` and `metric-store` this aborts startup, for the other sections it is
logged as an error and the section is applied only up to the offending key — so
watch the log after a configuration change.
Annotated examples are in [`configs/`](https://github.com/ClusterCockpit/cc-backend/tree/main/configs):
`config.json` (full featured, S3 archive, external metric stores),
`config-large.json` (large installation with LDAP/OIDC and DB tuning),
`config-demo.json` (minimal) and `uiConfig.json`.
Cluster topology and metric definitions are **not** part of `config.json`. They
are read from the `cluster.json` file inside each cluster's job archive
directory; `configs/cluster.json` is a documented example.
## Database Configuration ## Database Configuration
cc-backend uses SQLite as its database. For large installations, SQLite memory cc-backend uses SQLite as its database. For large installations, SQLite memory
@@ -272,8 +316,9 @@ The effective configuration is logged at startup for verification.
- [`cmd/cc-backend`](https://github.com/ClusterCockpit/cc-backend/tree/main/cmd/cc-backend) - [`cmd/cc-backend`](https://github.com/ClusterCockpit/cc-backend/tree/main/cmd/cc-backend)
contains the main application entry point and CLI implementation. contains the main application entry point and CLI implementation.
- [`configs/`](https://github.com/ClusterCockpit/cc-backend/tree/main/configs) - [`configs/`](https://github.com/ClusterCockpit/cc-backend/tree/main/configs)
contains documentation about configuration and command line options and required contains sample configuration files (`config.json`, `config-demo.json`,
environment variables. Sample configuration files are provided. `config-large.json`, `uiConfig.json`, `cluster.json`), example REST API
payloads, and the example rule sets for the job taggers.
- [`init/`](https://github.com/ClusterCockpit/cc-backend/tree/main/init) - [`init/`](https://github.com/ClusterCockpit/cc-backend/tree/main/init)
contains an example of setting up systemd for production use. contains an example of setting up systemd for production use.
- [`internal/`](https://github.com/ClusterCockpit/cc-backend/tree/main/internal) - [`internal/`](https://github.com/ClusterCockpit/cc-backend/tree/main/internal)
@@ -318,10 +363,14 @@ The effective configuration is logged at startup for verification.
- [`convert-pem-pubkey`](https://github.com/ClusterCockpit/cc-backend/tree/main/tools/convert-pem-pubkey) - [`convert-pem-pubkey`](https://github.com/ClusterCockpit/cc-backend/tree/main/tools/convert-pem-pubkey)
Tool to convert external pubkey for use in `cc-backend`. Tool to convert external pubkey for use in `cc-backend`.
- [`gen-keypair`](https://github.com/ClusterCockpit/cc-backend/tree/main/tools/gen-keypair) - [`gen-keypair`](https://github.com/ClusterCockpit/cc-backend/tree/main/tools/gen-keypair)
contains a small application to generate a compatible JWT keypair. You find contains a small application to generate a compatible Ed25519 JWT keypair for
documentation on how to use it the `auth.jwts` configuration.
[here](https://github.com/ClusterCockpit/cc-backend/blob/main/docs/JWT-Handling.md). - [`binaryCheckpointReader`](https://github.com/ClusterCockpit/cc-backend/tree/main/tools/binaryCheckpointReader)
- [`web/`](https://github.com/ClusterCockpit/cc-backend/tree/main/web) Dumps the contents of a metric store binary checkpoint file for inspection.
- `dataGenerator.sh` and `grepCCLog.pl`
Helper scripts to generate metric line protocol test data and to filter
cc-backend log output.
- [`web/`](https://github.com/ClusterCockpit/cc-backend/tree/main/web)
Server-side templates and frontend-related files: Server-side templates and frontend-related files:
- [`frontend`](https://github.com/ClusterCockpit/cc-backend/tree/main/web/frontend) - [`frontend`](https://github.com/ClusterCockpit/cc-backend/tree/main/web/frontend)
Svelte components and static assets for the frontend UI Svelte components and static assets for the frontend UI
+124 -6
View File
@@ -6,9 +6,20 @@ This release replaces the browser session implementation and requires a database
migration to version 12. Run `./cc-backend -migrate-db` after upgrading; the new migration to version 12. Run `./cc-backend -migrate-db` after upgrading; the new
`sessions` table is created automatically (a fresh `-init-db` also creates it). `sessions` table is created automatically (a fresh `-init-db` also creates it).
Existing login sessions are invalidated by the upgrade, so users have to log in Existing login sessions are invalidated by the upgrade, so users have to log in
again once. See the behavior changes below regarding the session cookie `Secure` again once.
flag and the removal of the `SESSION_KEY` environment variable.
For release specific notes visit the [ClusterCockpit Documentation](https://clusterockpit.org/docs/release/). **Two configuration changes are mandatory before starting 1.6.0:**
- The `.env` file is gone. Move every secret into the `auth` section of
`config.json` (or export the corresponding environment variable).
- The `main.resampling` section no longer accepts `trigger`, `resolutions` and
`minimum-points`. cc-backend rejects unknown keys inside `main` and aborts on
startup if they are still present. Replace them with `default-policy` and
`default-algo`.
See the behavior changes below for the details, as well as for the session
cookie `Secure` flag and the removal of the `SESSION_KEY` environment variable.
For release specific notes visit the [ClusterCockpit Documentation](https://clustercockpit.org/docs/release/).
## Changes in 1.6.0 ## Changes in 1.6.0
@@ -28,14 +39,117 @@ For release specific notes visit the [ClusterCockpit Documentation](https://clus
terminate TLS at a reverse proxy and want the `Secure` flag should serve terminate TLS at a reverse proxy and want the `Secure` flag should serve
cc-backend over HTTPS directly for now. cc-backend over HTTPS directly for now.
- **`SESSION_KEY` removed**: The `SESSION_KEY` environment variable is no longer - **`SESSION_KEY` removed**: The `SESSION_KEY` environment variable is no longer
used and should be removed from your `.env`. Server-side sessions use random used and can be dropped from your environment. Server-side sessions use random
tokens, so no cookie-signing secret is required. It is ignored if left in place. tokens, so no cookie-signing secret is required. It is ignored if left in place.
- **`.env` support removed (breaking)**: The `godotenv` bootstrap, the `.env`
file and `configs/env-template.txt` are gone. Secrets (JWT public/private key,
LDAP sync password, OIDC client id/secret, cross-login keys) now live in
`config.json` under `auth` next to the option that uses them. Each secret can
still be supplied through its existing environment variable
(`JWT_PUBLIC_KEY`, `JWT_PRIVATE_KEY`, `LDAP_ADMIN_PASSWORD`, `OID_CLIENT_ID`,
`OID_CLIENT_SECRET`, `CROSS_LOGIN_JWT_PUBLIC_KEY`, `CROSS_LOGIN_JWT_HS512_KEY`),
which takes precedence over the config value. `-init` now writes the demo JWT
keys into `config.json`.
- **Resampling configuration replaced by policies (breaking)**: The explicit
`trigger`, `resolutions` and `minimum-points` keys under `main.resampling` are
removed; `main` rejects unknown keys, so a config still carrying them aborts
startup. Configure `default-policy` (`low` ≈ 200, `medium` ≈ 500, `high` ≈ 1000
target points) and `default-algo` instead. Both are also exposed as per-user
settings, and the policy is now the single source for the requested resolution
and for the resampler's minimum-points threshold.
- **Default downsampling algorithm is now `average`**: Previously an unset
algorithm fell through to cc-lib's LTTB. Average consolidation makes every
plotted point the true arithmetic mean of its interval; LTTB deliberately keeps
extremes and cannot be read as interval means. Set
`main.resampling.default-algo` to `lttb` to restore the old behavior.
- **OIDC roles require an explicit mapping (breaking for role-based setups)**:
The hardcoded translation of token roles (which only recognised the literal
names `user`/`admin`/`manager`/`support` and silently dropped `api`) is
replaced by `auth.oidc.role-mapping`. The mapping is the sole source of roles:
unmapped token roles are ignored and users without a mapped role get the base
`user` role. Deployments whose IdP emits literal CC role names must now map
them explicitly (e.g. `"admin": "admin"`). Mapping targets are validated at
startup.
- **LDAP can assign elevated roles**: The new optional `auth.ldap.role-filters`
map (role → LDAP filter) grants `admin`/`support`/`api`/`manager` to accounts
matching a filter. LDAP is authoritative for the roles listed there — sync and
login both add *and remove* them to match group membership — while roles not
listed are preserved (a managed `manager` with assigned projects is never
stripped). Without `role-filters` the behavior is unchanged.
- **Unrecognized top-level config sections are reported**: cc-backend now logs a
warning for unknown top-level keys in `config.json`, which most often means a
setting was nested at the wrong level (e.g. `resampling` next to `main`
instead of inside it) and was silently ignored before.
- **The `auth` section is validated again**: Its JSON schema was missing the
enclosing `type`/`properties` wrapper, so it accepted every input and no auth
option was ever checked. With the schema fixed, an incomplete auth section now
aborts startup — `jwts` requires `max-age`, `oidc` requires `provider`, and
`ldap` requires `url`, `user-base`, `search-dn`, `user-bind` and `user-filter`.
Each of the three subsections stays optional on its own.
- **`archive.retention` target keys**: The retention target is configured with
`target-kind` and `target-path` (or the `target-*` S3 keys). The `location` key
used by the old example configuration was never read by the code, so a `move`
or `copy` policy configured that way ran against an empty target path.
### New features
- **Configurable footer links**: `main.footer-links` lets the Imprint and
Privacy Policy footer entries point at internal pages (default) or external
URLs. External `http(s)` targets open in a new tab; empty or unset values fall
back to the built-in `/imprint` and `/privacy` routes, and the existing
`./var/*.tmpl` override mechanism keeps working.
- **Display smoothing for metric plots**: A centered, NaN-aware moving average
can be applied in the plot layer, after whatever downsampling the backend
performed. It preserves series length and index alignment, skips NaN/null
samples, keeps all-NaN windows as gaps, and shrinks the window at the series
edges. The window is a per-user setting in data points with a site default in
`ui.plot-configuration.smoothing-window` (0 disables it). It is display-only:
reported statistics and job footprints are unaffected.
- **Metric tooltips**: `metricConfig` entries in `cluster.json` accept an
optional `tooltip` that is rendered in the metric selection dialog.
- **Metric store statistics from cached aggregates**: Buffers maintain running
sum/min/max/count while data is written and compute them eagerly on checkpoint
load, so a full-buffer statistics request is served from the cached aggregate
instead of rescanning the samples.
- **Checkpoint retention aware of running jobs**: The metric store receives a
`NodeProvider` (the job repository) at `Init`. Checkpoint cleanup now skips
nodes that still have running jobs, the full checkpoint history is loaded for
those nodes at startup, and this also applies to the `-cleanup-checkpoints`
CLI path.
### Bug fixes
- **Metric plots and the policy-based resample config**: `MetricPlot` still read
the removed `trigger`/`resolutions` values, so the array-based resolution
branch was unreachable and the zoom guard always fired. The zoom trigger is now
derived from the policy's target point count.
- **Resample algorithm for running jobs**: The selected algorithm was not
forwarded to the metric store when loading data for running jobs.
- **Partially covered buffers**: Requests that reach beyond the stored data now
get NaN padding at the start/end instead of misaligned series
(`dataNotAligned`).
- **Data race in the statistics read path**: Reading buffer statistics no longer
mutates the buffer, and the fast path is guarded against empty buffers.
- **Empty node levels leaked**: The retention free path now prunes node levels
that no longer hold any data, and no longer force-frees buffers that are
retained for running jobs.
- **Plot render options in the wrong place**: `PlotRenderOptions` is only shown
in the user settings again.
- **Invisible navbar**: Fixed a stuck animation that hid the navbar with certain
browser settings.
- **Subcluster dropdown**: The subcluster selection is size-limited and scrolls
instead of overflowing.
- **OIDC login session handling**: Fixed session handling on the OIDC login
route.
### Dependencies ### Dependencies
- **Added** `github.com/alexedwards/scs/v2` and - **Added** `github.com/alexedwards/scs/v2` and
`github.com/alexedwards/scs/sqlite3store`. `github.com/alexedwards/scs/sqlite3store`.
- **Removed** `github.com/gorilla/sessions` (and `github.com/gorilla/securecookie`). - **Removed** `github.com/gorilla/sessions` (and `github.com/gorilla/securecookie`),
and `github.com/joho/godotenv` together with the `.env` support.
- **Upgraded** `cc-lib` (GraphQL code regenerated) and refreshed the remaining Go
module dependencies.
## Changes in 1.5.4 ## Changes in 1.5.4
@@ -286,7 +400,11 @@ This is also the default.
- **cc-lib upgraded**: Updated to latest cc-lib version. - **cc-lib upgraded**: Updated to latest cc-lib version.
## Known issues ## Known issues (as of 1.5.2)
These were the open issues at the time of the 1.5.2 release and are kept here for
that release's record; they are not necessarily still open. Note that the
`ui-config` section referenced below is now called `ui`.
- The new dynamic memory management is not bullet proof yet across restarts. - The new dynamic memory management is not bullet proof yet across restarts.
Buffers that are kept outside the retention period may be lost across a Buffers that are kept outside the retention period may be lost across a