Upgrading TAS 5.17 to 5.19 – Technical Part (DevOps)

Technical guide for upgrading a customer TAS instance from v5.17 to v5.19. Intended for DevOps engineers and administrators who plan and carry out the upgrade: infrastructure, configuration, database migrations, deployment procedure and rollback.

Impacts on the customer configuration (templates, calculations, prints, workflow, variables) are covered in the follow-up article Upgrading TAS 5.17 to 5.19 – Consultant Part.
The upgrade contains migrations that irreversibly delete data (cron run history, invitations, Guides, instance graph). Do not start the upgrade without a full database backup – the migrations' down() restores only the schema, not the content.

TL;DR – what is most likely to trip you up

#

Risk

Impact

1

TAS_APP_VERSION is now mandatory

The backend will not start without it

2

A migration adds CHECK constraints on TEMPLATE_TASKS / TEMPLATE_GRAPH and NOT NULL + CHECK on JS_SCRIPTS.JS_TYPE without a backfill

The migration fails on legacy values

3

ArangoDB as a logger has been removed

Instances with TAS_LOGGER_ARANGO_ENABLE=true must switch to Elasticsearch before the upgrade

4

Certificates are migrated from the filesystem to secret_store

The directory must contain only certificates, otherwise junk gets ingested

5

Cron overhaul (tas-3500) – MSSQL table rebuild, CRON_RUNS dropped

Cron run history is lost; requires a window with no running cron workers

6

Zabbix switched from push to pull

The old integration stops delivering data

7

The secret_store.valid_from migration diverges on a 5.17 DB (5.17-only migration)

On PostgreSQL the type remains timestamp instead of timestamptz

8

Building images yourself? plugin/ is no longer in the backend build context; the frontend has new build target names

The plugin build fails / the wrong frontend target gets deployed

9

security.password.saltRounds / pepper moved under bcrypt, encryptionIvLength can no longer be a string

Configuration validation rejects the old local override

Checklist

Complete this before you touch anything else.

Migration data prerequisites

Run against a copy of the production DB. Every row in the result means the migration will fail.

(A) JS_SCRIPTS.JS_TYPE – newly NOT NULL + check in ('C','F','R'), no backfill. The legacy default was 'B'.

SELECT "JS_ID", "JS_TYPE" FROM "JS_SCRIPTS"
WHERE "JS_TYPE" IS NULL OR "JS_TYPE" NOT IN ('C','F','R');

(B) TEMPLATE_TASKS – new CHECK constraints.

SELECT DISTINCT "TTASK_TYPE"                FROM "TEMPLATE_TASKS";  -- allowed: A,S,P,E,N,W,C
SELECT DISTINCT "TTASK_ASSESMENT_METHOD" FROM "TEMPLATE_TASKS"; -- S,U,T,L,W,C,A,V,P
SELECT DISTINCT "TTASK_ASSESMENT_HIERARCHY" FROM "TEMPLATE_TASKS"; -- G,C,D,P,A,S,L
SELECT DISTINCT "TTASK_INVOKE_EVENT" FROM "TEMPLATE_TASKS"; -- I,B
SELECT DISTINCT "TTASK_ENOT_TGT_TYPE" FROM "TEMPLATE_TASKS"; -- O,G,T,P,S,U,R
-- same for TTASK_ENOT_COPY_TYPE, TTASK_ENOT_BLIND_TYPE, TTASK_ENOT_REPLY_TYPE
SELECT DISTINCT "TTASK_CAN_BE_BUTTON", "TTASK_GEN_HISTORY",
"TTASK_IS_BULK_COMPLETABLE", "TTASK_MULTIINSTANCE_FLAG",
"TTASK_SUFFICIENT_END" FROM "TEMPLATE_TASKS"; -- Y,N only

(C) TEMPLATE_GRAPH

SELECT DISTINCT "TGRAPH_COLOR"     FROM "TEMPLATE_GRAPH";
-- allowed: fffacd, e6e6fa, ffe4e1, f0ffff, ffe4b5, ffffff, ''
-- (the migration lowercases the values first)
SELECT DISTINCT "TGRAPH_BPMN_TYPE" FROM "TEMPLATE_GRAPH";
-- allowed: bpmn:UserTask, bpmn:SendTask, bpmn:ReceiveTask, bpmn:ServiceTask,
-- bpmn:SubProcess, bpmn:IntermediateThrowEvent, bpmn:IntermediateCatchEvent,
-- bpmn:ScriptTask, bpmn:Task, bpmn:Participant, bpmn:Lane

(D) Orphan processes – the migration adds an FK INSTANCE_PROCESSES → TEMPLATE_PROCESSES. It sets TPROC_ID = NULL on orphans by itself; determine the extent in advance.

SELECT count(*) FROM "INSTANCE_PROCESSES" ip
WHERE ip."TPROC_ID" IS NOT NULL
AND NOT EXISTS (SELECT 1 FROM "TEMPLATE_PROCESSES" tp WHERE tp."TPROC_ID" = ip."TPROC_ID");
-- same for ARCH_INSTANCE_PROCESSES

(E) Duration estimate of the heaviest steps

SELECT count(*) FROM "TEMPLATE_TASK_LINKS";            -- migration #18 performs row-by-row UPDATEs
SELECT count(*) FROM "INSTANCE_PROCESS_HISTORY"; -- migration #22 casts date -> timestamp
SELECT count(*) FROM "ARCH_INSTANCE_PROCESS_HISTORY"; -- on MSSQL via an intermediate nvarchar(max) step
SELECT count(*) FROM "INSTANCE_TASKS"; -- column drops + CHECK validation
SELECT count(*) FROM "ARCH_INSTANCE_TASKS";
Values outside the allowed set must be fixed in the data before the upgrade. Fixing them is the consultant's job – hand over the query results and follow the consultant part. The alternative is modifying the migration locally, but that means deviating from the standard branch – always document such an intervention.

Environment blockers

Check

Expected state

TAS_LOGGER_ARANGO_ENABLE

must be false / unused; the ArangoDB logger has been removed

Elasticsearch

mandatory for logs, cron run history and DMS fulltext

Redis

practically mandatory (fleet health, cron monitor, plugin restart, CRON_LAST_RUN); TAS_CACHE_MODE=memory is for single-instance dev only

Does the customer use the ksef plugin?

Blocking – it does not exist in 5.19; resolve with development before the upgrade

Does the customer use the old EWS mail crons?

The migration deletes them; their configuration can be found in the update log under the [tas-3432] tag

Custom image build?

--build-arg TAS_APP_VERSION=<tag> is mandatory; additionally --build-context plugin_source=./backend/plugin and the new frontend target names

Local override config/local.*?

Check security.password.* (moved under bcrypt) and database.encryptionIvLength (can no longer be a string)

Elasticsearch capacity

A new index <TAS_ELASTIC_INDEX_NAME>-cron-runs will be added – plan its retention

Plugin volume

Must be writable at runtime and shared by all instances (plugin installation from the GUI)

Zabbix monitoring

Prepare an API token (admin scope) and the new template

MSSQL on a schema other than dbo

Fixed in tas-3532 (also included in 5.17.9), but verify that you are upgrading from version ≥ 5.17.9

MSSQL version

Migrations use IF EXISTS → SQL Server 2016+

Backup

Migrations irreversibly delete data. A full DB backup is required – the migrations' down() restores only the schema, not the content.

Objects removed by the upgrade:

  • CRON_RUNS – cron run history (moved to Elasticsearch, old data is not transferred),
  • GUIDES, HEALTH_STATUS_PARAM,
  • INSTANCE_TASK_INVITATIONS, TEMPLATE_TASK_INVITATIONS,
  • INSTANCE_TASK_LINKS, INSTANCE_LINK_CONDITIONS,
  • INSTANCE_GRAPH, ARCH_INSTANCE_GRAPH,
  • TEMPLATE_TASK_CALCULATIONS,
  • columns: ITASK_SEEN, ITASK_DISC_FLAG, TTASK_DISC_FLAG, TTASK_ENOT_BODY, TTASK_IS_PROC_ID, TTJSCALC_TITLE, ITJSCALC_TITLE, TTASKLINK_IS_MANDATORY, IGRAPH_ITASKCON_ID, CRON_ALIAS, CRON_IS_CLONE, CRON_TIMEOUT, CRON_HELP.

Infrastructure

Runtime and dependencies

Neither package.json has an engines field – the Node version is determined solely by the base image.

Item

5.17

5.19

Node.js (backend image)

24.13.0

24.16.0 (backend/.nvmrc, node:24.16.0-alpine)

Node.js (frontend image)

24.11.0

24.11.0 – unchanged

TypeScript

5.9.3

6.0.3

Babel core/preset-env

7.29

8.0.x (ESM-only)

Fastify

5.8.1

5.10.0

@fastify/multipart

9.4.0

10.0.0

MikroORM

6.6.8

6.6.15 (deliberately not v7)

tedious (MSSQL)

19.2.1

19.2.1 – stays on the version pinned by @mikro-orm/mssql@6

pg

8.20.0

8.22.0

@elastic/elasticsearch

9.3.3

9.4.2

ioredis / bullmq

5.10.0 / 5.70.2

5.11.1 / 5.79.3

nodemailer

8.0.1

9.0.3

firebase-admin

13.7.0

14.1.0 (modular API)

puppeteer

24.38.0

25.3.0

sharp

0.34.5

0.35.3

ejs

4.0.1

6.0.1

htmlparser2 / node-html-parser

10.1.0 / 7.1.0

12.0.0 / 9.0.0

csv-parse / js-beautify / ini / basic-ftp

6.1.0 / 1.15.4 / 6.0.0 / 5.2.0

7.0.1 / 2.0.3 / 7.0.0 / 6.0.1

rate-limiter-flexible

9.1.1

11.2.0

npm in the backend image

default

npm@latest (v12) + patched bundled deps

New backend dependencies: @modelcontextprotocol/sdk, @fastify/sse, fastify-plugin, tar-stream, tmp, tsx (dev).

Removed: arangojs, typedi (replaced by a custom DI in backend/src/infrastructure/di/service.ts), reflect-metadata, systeminformation, docx-templates, docxtemplater, pizzip (moved to the docxGenerator plugin).

Frontend: ESLint + Prettier → Biome 2.4.x; added Vitest 4 + Testing Library + jsdom; react-beautiful-dnd → @dnd-kit; removed hopscotch (discontinued Guides) and npm as a frontend dependency. Security bumps: axios 1.16.1, sanitize-html 2.17.4, dompurify 3.4.5, express 4.22.2, lodash 4.18.1, overview-components 1.1.184. React / MUI / webpack major versions unchanged.

Dockerfile

The docker/** directory (swarm stacks tasdev.yml, tasmssql.yml, tasmssql-mac.yml, tasoracle.yml, scripts) is unchanged between 5.17 and 5.19. No new services, volumes or healthchecks. There is no docker-compose*.yml in the repo, and neither Dockerfile contains a HEALTHCHECK instruction.

Backend (backend/Dockerfile)

  • ARG TAS_APP_VERSION → ENV TAS_APP_VERSION baked into every stage. CI passes it from the git tag (TAG_NAME_CLEANED, including the prerelease suffix).
  • npm ci --legacy-peer-deps (because of the openapi-typescript peer dep), in both the prod and dev stages.
  • npm_config_allow_remote=root, npm_config_allow_git=root – npm 12 would otherwise reject the xlsx CDN tarball and the annotpdf git dependency.
  • The plugin source is no longer in the backend build context (backend/.dockerignore); plugin_builder receives it via a named build context. plugin_builder now inherits FROM test (previously FROM linter).
  • Unit tests are no longer run during the build (they run in a separate CI job).
  • The set of apk packages (LibreOffice, OpenJDK17, Chromium, msttcorefonts, openssl) is unchanged.
  • Production target names are unchanged: prod, prod_obfuscated (default), prod_devlicense, prod_obfuscated_devlicense.

If you build the image manually, you must use:

docker buildx build --build-context plugin_source=./backend/plugin --target plugin_archives ...
Without --build-context plugin_source the plugin build fails.

Frontend (frontend/Dockerfile) – check your deployment pipeline

The build has been split by license. 5.17 had source → prod / prod_obfuscated. 5.19 has:

source_base → source_prodlicense → prod, source_obfuscated_prodlicense → prod_obfuscated (default)
→ source_devlicense → prod_devlicense, source_obfuscated_devlicense → prod_obfuscated_devlicencse

The build patches developmentBuild in frontend/src/components5.0/zustand/loggedUserStore.ts according to the license.

The name of the dev-license obfuscated target contains a typo: prod_obfuscated_devlicencse. If you target it, you must reproduce the typo.

New and changed env variables

Newly mandatory

Variable

Note

TAS_APP_VERSION

Previously optional with a fallback of "5" from manifest.json. Now getMandatoryEnvString – the backend crashes on startup. manifest.json and readManifestVersion() have been removed. The value propagates to application.version, GET /plugins/version, GET /system-health and to the plugin compatibility check.

Removed

Variable

Replacement

TAS_LOGGER_ARANGO_ENABLE, TAS_LOGGER_ARANGO_URL, TAS_ARANGO_DB_NAME, TAS_ARANGO_USERNAME, TAS_ARANGO_PASSWORD, TAS_ARANGO_JWT

Elasticsearch (loggerType is always elastic)

TAS_LOGGER_ELASTIC_ENABLE

no longer relevant

TAS_CERTIFICATES_STORAGE_PATH (from the configuration declaration)

Certificates are stored in secret_store / Vault. However, the migration still reads it via process.env (default /app/tas/storage/certificates) – see Certificate migration to Vault.

New (all optional)

Group

Variables

Cache

TAS_CACHE_MODE (redis | memory, default redis) – a new fallback option; in 5.17 there was no alternative to Redis

Crons

TAS_CRONS_CONSECUTIVE_FAILURE_THRESHOLD (actual default 5; when set, it is read-only in the GUI)

Elastic

TAS_ELASTIC_CRON_RUN_INDEX_NAME (default <TAS_ELASTIC_INDEX_NAME>-cron-runs)

System health

TAS_HEALTH_SAMPLE_INTERVAL_SEC (60), TAS_HEALTH_RETENTION_HOURS (48), TAS_HEALTH_STALE_AFTER_SEC (180)

SIEM – file

TAS_LOGGER_FILE_ENABLE, TAS_LOGGER_SIEM_FILE_PATH (default /app/tas/storage/log/tas.log), TAS_LOGGER_FILE_LOG_LEVEL (default 3 = INFO)

SIEM – syslog

TAS_LOGGER_SYSLOG_ENABLE, TAS_LOGGER_SYSLOG_HOST, TAS_LOGGER_SYSLOG_PORT (514), TAS_LOGGER_SYSLOG_PROTOCOL (udp | tcp), TAS_LOGGER_SYSLOG_FACILITY (1), TAS_LOGGER_SYSLOG_LOG_LEVEL (2 = WARN+)

Console fallback

TAS_LOGGER_CONSOLE_FALLBACK_ENABLE (default true) – with loggerType=elastic and stdout disabled, a console handler is added automatically so that logs are not lost during an ES outage

Mail via MS Graph

TAS_MAIL_SMTP_TYPE=msgraph + TAS_MAIL_MSGRAPH_TENANT_ID, _CLIENT_ID, _CLIENT_SECRET, _SENDER, _SCOPE

Mail

TAS_MAIL_ALLOWED_CUSTOM_HEADERS (allowlist of custom X-headers)

Passwords

TAS_SECURITY_ARGON2_MEMORY_KIB (47104), _ITERATIONS (1), _PARALLELISM (1), _HASH_LENGTH (64), _SALT_LENGTH (32)

Vault

TAS_VAULT_KEY_PAIR_PUBLIC_SUFFIX (.public), TAS_VAULT_KEY_PAIR_PRIVATE_SUFFIX (.private)

envVariables.ts documents a default of 3 for TAS_CRONS_CONSECUTIVE_FAILURE_THRESHOLD, but staticConfigHandler.ts uses 5. The code wins – the actual default is 5.

Dead, but still declared

TAS_LOGGER_FILE_PATH remains in envVariables.ts, but staticConfigHandler.ts no longer reads it (logging.fallbackfilePath has disappeared from globalConfig.ts). It has been replaced by TAS_LOGGER_SIEM_FILE_PATH.

Stricter validation of the static configuration – it may reject a local override that has worked so far.

Key

5.17

5.19

database.encryptionIvLength

number | string

number only

security.password.saltRounds, security.password.pepper

flat keys

moved under security.password.bcrypt.{saltRounds,pepper}

–

–

new mandatory blocks security.password.argon2id.* and security.password.currentHashAlgorithm (bcrypt | argon2id)

logging.file.*, logging.syslog.*, application.health.*, featureFlag.vault.{keyPairPublicSuffix,keyPairPrivateSuffix}, integrations.mail.allowedCustomHeaders, integrations.elastic.cronRunIndex

–

newly mandatory in the schema (defaults exist, so they pass without any intervention)

cache

required: ["redis"]

required: ["mode"]

Hex validation of TAS_DATABASE_ENCRYPTION_KEY (64 hex characters) and of token secrets (hex, min. 64 characters) is the same in both branches – it was already introduced in 5.17 with tas-3079. It is not a new obstacle.

Frontend

Variable

Meaning

TAS_PLUGIN_STORE_URL

Base URL of the plugin store (frontend/config/tas.base.js → config.pluginStoreUrl). When it is not set, the “Plugin Store” tab is hidden. The browser calls the store directly; the backend does not access it. It is not in docker/devstack/.env.template – add it manually.

Dynamic (GUI) configuration

Change

Detail

New

application.crons.consecutiveFailureThreshold (default 5)

New

application.health.* (sample interval / retention / stale threshold)

New

mail.sendMail.{newline,path} and mail.msgraph.{tenantId,clientId,clientSecret,sender,scope}; the mail.type enum has been extended with sendmail and msgraph

New

ai.recursionLimit (LangGraph), plus AI enabled/provider/key/Azure OpenAI fields

New

featureFlag.gui.disableMobilePullToRefresh

Renamed

security.password.expirationEmail → expirationEmailEnabled (the mapping is in dynamicConfigMigrationConsts.ts and is applied automatically)

Removed

storage.certificates (also dropped from required)

New tab

Administration → Configuration → Plugins (plugin configuration in PLUGIN_CONFIG, handler pluginConfigHandler.ts)

Static only

integrations.ai.mcpServers (definitions of outgoing MCP servers) – not editable in the GUI

Dynamic configuration changes can also be applied from the CLI: npm run dynamic-config:update.

Database migrations

24 new logical migrations, each as an MSSQL + PostgreSQL pair (none of them is dialect-neutral). Plus regenerated snapshots (.snapshot-TAS.json for both dialects).

Migration overview

#

Migration

What it does

Risk

1

Migration20260420170448

CREATE TABLE AI_CONVERSATION_MEMORY

low

2

Migration20260424112724/25

secret_store + valid_from, kind check extended with certificate

see valid_from divergence

3

…-vault-certificate-migration

TS data migration – reads files from the certificate directory, encrypts them and inserts them into secret_store

high

4

Migration20260424130139

USERS.USER_PASSWORD 60 → 255 characters (for argon2id)

low (table rewrite on PG)

5

…-ai-llm-tool-call-status-fix

check constraint on AI_LLM_TOOL_CALL.STATUS + pending_confirmation

low

6

…-remove-backend-app-id-seq

DROP SEQUENCE BACKEND_APPLICATION_ID_SEQ

low

7

Migration20260525142332/33

AI_MESSAGE.CLIENT_ACTION

low

8

Migration20260527075819

CREATE TABLE PLUGIN_CONFIG

low

9

Migration20260529132931/32

*_ENOT_CUSTOM_HEADERS added to 3 notification tables

low

10

Migration20260602145822/23

DROP INSTANCE_TASK_INVITATIONS, TEMPLATE_TASK_INVITATIONS

data loss

11

…-remove-guides

DROP GUIDES + GUIDE_ID_SEQ

data loss

12

Migration20260608192313/14

Major template cleanup: drop TEMPLATE_TASK_CALCULATIONS, drop 6 columns (incl. on INSTANCE_TASKS and ARCH_INSTANCE_TASKS), type changes, lowercase TGRAPH_COLOR, ~15 new CHECK constraints

high + long-running

13

Migration20260609183923/24

DROP COLUMN ITASK_SEEN on INSTANCE_TASKS + ARCH_

medium (large tables)

14

…-cron-overhaul

tas-3500. MSSQL: rebuild of the CRONS table (CRONS_NEW + IDENTITY_INSERT + drop + rename), drop FK from XML_PROCESS_IMPORT, DROP CRON_RUNS. PG: in-place column renames, backfill alias→name, CRON_STATUS→boolean IS_ACTIVE, TYPE NOT NULL + check, dedupe + UNIQUE on DISPLAY_NAME

highest

15

…-add-markdown-webp-file-types

seeds text/markdown, image/webp into DMS_FILE_TYPE (idempotent)

low

16

Migration20260626183639

DROP INSTANCE_TASK_LINKS, INSTANCE_LINK_CONDITIONS + sequences + ROLE_ACCESS_RIGHTS rows; drop IGRAPH_ITASKCON_ID, TTASKLINK_IS_MANDATORY

high

17

…-drop-health-status-param

DROP HEALTH_STATUS_PARAM

low volume

18

…-template-link-conditions-tree

The largest data migration. TTASK_OUTGOING_SPLIT_MODE NOT NULL default 'I'; CONDITIONS jsonb + TTASKLINK_IS_DEFAULT + UPDATED_AT/UPDATED_BY on TEMPLATE_TASK_LINKS; loads all links + TEMPLATE_LINK_CONDITIONS + TEMPLATE_VARIABLES into memory, builds the condition tree, row-by-row UPDATE, backfill, NOT NULL, count verification with a throw on mismatch, sequence fix, new FKs

high + long-running

19

…-instance-process-tproc-fk

sets TPROC_ID = NULL on orphans, then FK → TEMPLATE_PROCESSES

long-running (largest tables)

20

…-prnt-append-scripts

PRNT_APPEND_SCRIPTS, TTASK_APPEND_SCRIPTS; JS_SCRIPTS.JS_TYPE → NOT NULL + check ('C','F','R') without backfill

high

21

…-drop-instance-graph

DROP INSTANCE_GRAPH, ARCH_INSTANCE_GRAPH, IGRAPH_ID_SEQ, ROLE_ACCESS_RIGHTS rows; new FK ARCH_INSTANCE_TASK_HISTORY → ARCH_INSTANCE_TASKS

high

22

…-iph-inserted-datetime

IPH_INSERTED date → datetime2/timestamptz on INSTANCE_PROCESS_HISTORY + ARCH_ (MSSQL via an intermediate nvarchar(max) step)

long-running – full rewrite of large history tables

23

…-mcp-tool-call-audit

AI_LLM_TOOL_CALL + USER_ID (FK), COLLECTION, SOURCE; backfill 'unknown'/'general', then NOT NULL

medium

24

…-seed-default-crons

seeds 9 default crons – only if CRONS is empty

low

Existing instances have CRONS populated, so migration #24 is a no-op. New crons (DatabaseHealthReportCron, DatabaseIndexRebuildCron, …) must be created manually via Administration → Crons → “+”.

Certificate migration to Vault (#3)

The migration walks through the directory process.env.TAS_CERTIFICATES_STORAGE_PATH (fallback /app/tas/storage/certificates), parses every file in it as X.509, encrypts it and inserts it as a secret_store row with kind = certificate.

Before running the migration, verify:

  • that the variable points to the correct directory and that the directory is available in the migration's runtime environment (volume mounted),
  • that the directory contains only certificates – the migration will try to import anything else (README files, keys, temp files),
  • that globalThis.container (database encryption key) is initialized and that a user with id = 1 exists.
The class in the file is named Migration20260425100000, while the file is Migration20260424112725/26-vault-certificate-migration.ts. This has no functional impact, but when searching manually in the migration table, do not go by the file name.

secret_store.valid_from divergence between 5.17 and 5.19

5.17 added the column via the migration Migration20260528120000-secret-store-valid-from. This migration does not exist in 5.19 – the column is added by the earlier Migration20260424112724 (MSSQL) / …725 (PG).

Consequences for a DB that has already gone through 5.17:

  • An orphaned record Migration20260528120000 remains in the migration table.
  • Migration20260424112724/25 will run “out of order”. Both are written defensively (MSSQL if not exists (select 1 from sys.columns …), PG add column if not exists), so they will not fail – this is the tas-3673 fix; previously this ended with MSSQL error 2705.
  • PostgreSQL – type mismatch. 5.17 created valid_from as timestamp, whereas 5.19 would create it as timestamptz. The migration skips the column, so it stays timestamp and does not match the 5.19 snapshot/entity. Functionally this does not matter for now, but the next migration generated from the snapshot may try to address it. Verify after the upgrade and align manually if needed.
  • breakpointCheck (npm run mig) may flag the orphaned record – use npm run mig:nocheck or the new tools below.

New CLI for troubleshooting migrations

npm run db:migrations                 # JSON dump of the migration table: name, time, what is pending
npm run db:migrations:set-state -- name=<Migration…> state=executed # marks it as executed (the SQL is not run)
npm run db:migrations:set-state -- name=<Migration…> state=pending # deletes the row → it will run again
Neither direction touches the schema. state=executed assumes that the schema already is in the target state; state=pending causes the SQL to run again against a schema that may already contain its changes. The output includes willReRunOnNextMigration (for a deleted row without a matching file it stays false).

Other new commands: npm run db:report (index fragmentation report + heavy queries), npm run db:rebuild-indexes.

Post-migration step outside the migration table

Every migrate up now runs cron parameter normalization against the parametersSchema of the respective cron class:

  • type mismatches are examined and coerced ("5" → 5),
  • properties not allowed by the schema are removed,
  • defaults are not filled in,
  • a row that still does not validate after normalization (or could not be parsed) is left untouched and a warning is logged.
After the migration, go through the log and look for the recovery records (they contain both the original and the normalized JSON). Unfixable rows are logged again on every migration until someone fixes them in the administration.

Recommended upgrade procedure

  1. Announce the downtime. Estimate the window using query (E) – migrations #14, #18, #19 and #22 scale linearly with the number of rows.
  2. FULL DB BACKUP + snapshot of the storage volume (certificates, plugins, DMS).
  3. Pre-flight check of the data prerequisites and environment blockers. Fix data defects NOW.
  4. Stop operations:
    • first the cron workers (required – the cron overhaul drops columns in a single phase),
    • then the web backends,
    • put the frontend / LB into maintenance.
  5. Prepare the configuration for 5.19:
    • TAS_APP_VERSION (mandatory),
    • remove TAS_LOGGER_ARANGO_* and TAS_LOGGER_ELASTIC_ENABLE,
    • verify TAS_CERTIFICATES_STORAGE_PATH for the migration,
    • optionally TAS_PLUGIN_STORE_URL on the frontend.
  6. Deploy the 5.19 images (backend + frontend), but DO NOT START the backends yet.
  7. Run the migrations from a single instance: npm run mig. Watch the log; if it fails, do not retry blindly – see the new CLI.
  8. Check the migration log:
    • warnings from cron parameter normalization,
    • lines tagged [tas-3432] = configuration of the deleted EWS crons → save it aside,
    • the result of the certificate migration into secret_store.
  9. Deploy / update the plugins to the 5.19 versions. Without this they will not load.
  10. Start the web backends → verify /status/liveness, /status/readiness, /status/startup.
  11. Start the cron workers.
  12. Perform the post-upgrade configuration.
  13. Go through the verification checklist.
  14. Turn off maintenance mode.
The probes for Kubernetes/LB (/status, /status/db, /status/db/users, /status/liveness|readiness|startup) are unchanged.

Post-upgrade configuration

Plugins – mandatory

Manifests in 5.19 have minimalCompatibleTasVersion: "5.19.0" and maximalCompatibleTasVersion: "5.19.999". Because application.version now contains the actual version (previously it was masked by an outdated manifest.json with 5.18.16), an incompatible plugin is silently skipped.

Plugin versions in 5.19:

Plugin

sysname

Version

Ollama

Ollama

1.0.6

API Extensions

apiExtensions

1.0.4

Azure OpenAI

AzureOpenAI

1.0.7

Claude AI

ClaudeAI

1.0.7

docx Generator (new)

docx-generator

1.0.2

Expose DB

expose-db

1.0.7

Google GenAI

GoogleGenAI

1.0.7

Help Docs (new)

helpDocs

1.0.5

isDoc

is-doc

1.0.4

OpenAI

OpenAI

1.0.8

PDF Annotation

pdf-annotation

1.0.4

Single Sign-On

single-sign-on

1.0.3

Template Validation

template-validation

1.0.10

The ksef plugin is missing in 5.19 (it was present in 5.17). If the customer uses it, escalate to development before the upgrade.

New plugin management options:

  • Administration → Plugins → Plugin Store tab (only when TAS_PLUGIN_STORE_URL is set),
  • POST /plugins/install – uploads the archive to the plugin volume; this does not load the plugin,
  • DELETE /plugins/:sysname – uninstall (deletes the folder, restartRequired: true),
  • POST /plugins/kill – broadcast via Redis; each backend (web and cron) sends SIGTERM to its own graceful-shutdown handler and the orchestrator starts it again,
  • a plugin with a broken manifest or a capability that cannot be loaded no longer crashes backend startup – it is logged and skipped (tas-3545).
POST /plugins/kill requires the deployment to restart terminated containers (docker / k8s restart policy). Under a bare npm run dev, the process simply exits.

Crons – requires attention

Complete overhaul (tas-3500). What to do after the upgrade:

  1. Go through the entire list of crons in Administration → Crons. The column renames (CRON_NAME→DISPLAY_NAME, CRON_FILE→TECHNICAL_NAME, CRON_SYNTAX→TIMING, CRON_STATUS→boolean IS_ACTIVE) have been applied to the data, but verify that the right crons are active.
  2. Run history has been lost – CRON_RUNS is dropped; new history goes to the Elasticsearch index <integrations.elastic.index>-cron-runs (created automatically via iniESIndex; no-op when ES indexing is disabled).
  3. Cloned crons no longer exist (CRON_ALIAS, CRON_IS_CLONE dropped). The alias has been migrated into the user-facing name.
  4. The per-cron timeout no longer exists (CRON_TIMEOUT), and neither does the POST /cron-runs/kill endpoint or the “kill run” action. Crash recovery via stale heartbeat and the “running much longer than usual” indicator remain.
  5. Cron help no longer exists – parameter documentation is now in parametersSchema and is rendered by JsonForms.
  6. Set application.crons.consecutiveFailureThreshold (default 5) – after N consecutive failures, a notification e-mail is sent to the cron's error addresses.
  7. Consider activating the new crons (they are created manually and are inactive by default): DatabaseHealthReportCron, DatabaseIndexRebuildCron.
  8. The DeleteLogs cron has a new default name and description (“Remove old logs.” – the word “arango” has been dropped).
  9. “Factory settings” on the cron detail now resets unconditionally (timing, description, parameters and active state; the user-facing name is kept).
  10. Crons with parameters that do not match the schema can no longer be saved – POST /crons/create and POST /crons return 400 INVALID_CRON_PARAMETERS.
  11. Behavior change: MailReportsCron treats a task as overdue when deadline < now (previously < today's midnight).

Zabbix – push → pull, reconfiguration required

Area

5.17

5.19

Model

TAS pushes metrics (DataSender/ZabbixService timer)

Zabbix scrapes

Configuration

HEALTH_STATUS_PARAM table + /health-status/settings*

removed

Template

backend/src/api/health/zabbix_default.json

backend/src/service/health/zabbix/tas_template.yaml (Zabbix 6.0 HTTP agent)

Endpoint

POST /health-status

GET /system-health/metrics

Auth

basic auth hook

long-lived B2B API token (admin scope), Authorization: Bearer header, macro {$TAS.HEALTH.TOKEN}

Steps: generate an API token with admin scope → import tas_template.yaml → set the macro → verify LLD discovery. The scrape endpoint deliberately bypasses the isAlive/inMaintenance hooks, so monitoring keeps working even during draining/maintenance.

The keys tas.crons and tas.cronStatus and the /health-status crons actions (cronStatus/crons/cronsGui/cronsHealth) have kept their field shape – existing monitoring based on them keeps working. Removed: scheduledTasksGui, the tasZombieCron metric.

The maxDiskUsePct metric and the per-disk details from the system probe are gone (the systeminformation dependency was removed; node:os is read instead). memUsedPct is now slightly higher and more conservative (os.freemem()).

Appstatus page removed, new System Health page

The legacy Appstatus page is gone. It is replaced by Administration → System Health (/administration/system-health) with probes for: application, database, database-activity, Elasticsearch, Redis, Tika, LibreOffice, e-mail, Firebase, crons, host resources and business KPIs.

  • Fleet-aware: each instance is sampled every 60 s into Redis (HEALTH_FLEET hash + per-instance series, 48h window / 49h TTL). Without Redis it degrades to the local instance only.
  • Database Activity reads pg_stat_activity / sys.dm_exec_*. The MSSQL user needs VIEW SERVER STATE, otherwise the card is shown as “unknown” (previously it would print the entire failing SQL).
  • New dashboard widget “System health” (admin only).

E-mail

  • New msgraph transport (app-only / Client Credentials) alongside basic, oauth2 and sendmail.
  • Custom X-headers in notifications with placeholder support + allowlist (TAS_MAIL_ALLOWED_CUSTOM_HEADERS).
  • MS Graph mail crons: clientSecret can be a {{vault:name}} reference to a Vault secret; new “Discover folders” button (POST /ms-graph/discover-folders) and mailbox duplication.
  • Newly created “Create Processes from MS Graph mails” crons have useEmailObjectInDataHolder = false (so that field mapping works). Existing crons keep their stored parameters.

Passwords and authentication

  • Hashing has switched to argon2id with lazy migration from bcrypt – a user's password is rehashed on their next login. The global pepper has been removed, including the related config.
  • USERS.USER_PASSWORD extended to 255 characters.
  • Fallback to another authority on the default login path (without an explicit auth_id).
  • Locking a user now invalidates the cache (USER2-<id>) → existing access tokens stop working immediately.
  • GET /authorization/config returns paramsSchema per authority (the frontend renders forms from the JSON schema).

Breaking changes for integrations and API

Removed endpoints

Endpoint

Note

POST /health-status, GET /status/full, /health-status/settings*, test/default-json

replaced by GET /system-health*

GET /connections/used-connections, POST /connections/destroy-connection

replaced by GET /system-health/db-activity + POST …/cancel

POST /ai/chat

/ai/messages remains

GET /guides/:id?, POST /guide/:id?, DELETE /guide/:id

the Guides feature has been discontinued

GET user password status

removed without replacement

POST /crons/clone, POST /crons/restart-cron, GET /cron-runs/summary, POST /cron-runs/kill

removed

POST /tasks/invitation/:itaskId

the “invitation” task type has been discontinued

GET /template-link/:id

links are read via the new model

Changed contracts

What

Change

POST /processes/to-delete/:iprocId, /processes/suspend/:iprocId, /processes/resume/:iprocId

consolidated into POST /processes/:iprocId/status

GET /roles/:id?

split into a list + GET /roles/:id

GET /template-processes/:tProcId/template-tasks/:id

ttask_event_params is returned as a raw JSON string – the client must JSON.parse it. Previously the keys were recursively lowercased, which broke camelCase inside the smart-event configuration.

GET /processes/:id/graph

Payload rewritten: one node per template task with instances / history_entries fields, a representative solver, drift markers origin (both/instanceOnly/templateOnly). Now enforces process rights – an inaccessible case returns 400 LACK_OF_PERMISSIONS instead of an empty diagram.

/crons*

renamed fields: cronTechnicalName, DISPLAY_NAME, TIMING, PARAMETERS, TYPE, boolean isActive; the schema is parametersSchema (previously cronParametersSchema in some places)

GET /org-units/import-preview

now returns only headers + checkedEntities (tree visualization removed)

Error states from 3rd-party libraries

An error from a library with its own statusCode (typically Elasticsearch 401) is no longer propagated – 500 is returned instead. The status is determined only by TAS exceptions and Fastify errors (FST_*). Genuine auth errors (AuthException) keep working.

MCP tool names

changed from CollectionName__tool to a bare tool; the collection is in the tool description as [CollectionName] …

New endpoints

  • POST /mcp/:scenario? – stateless MCP Streamable HTTP endpoint for external clients (Claude Desktop / Claude Code). Scenarios: case-creation, case-editing, case-lookup, knowledge, account, administration; a bare /mcp exposes everything. API token only (tokenPayload.kind === "api"); a session token returns UNAUTHORIZED. Every call is audited into AI_LLM_TOOL_CALL with source = mcp.
  • GET /system-health, /system-health/series, /system-health/db-activity, POST /system-health/db-activity/cancel, GET /system-health/metrics
  • GET /crons/available, GET /crons/monitor, POST /crons/create, /cron-runs, /cron-runs/:id
  • GET /plugins/version, POST /plugins/install, POST /plugins/kill, DELETE /plugins/:sysname
  • GET /plugin-config, /plugin-config/:sysname/schema, /plugin-config/:sysname/data, POST /plugin-config
  • GET /scripts/list/:type, GET /scripts/mapped/:kind/:kindId?
  • POST /ms-graph/discover-folders
  • DELETE license, vault key-pair generation, AI observability endpoints
MCP limitation: write tools that require confirmation in the chat are executed immediately via MCP – the protocol does not enforce a server-side human-in-the-loop (destructiveHint is advisory only).

Post-upgrade verification checklist

Startup and basics
  • The backend started without "ENV variable … is mandatory" in the log
  • GET /plugins/version returns the actual version (5.19.x), not 5.18.16
  • /status/liveness, /status/readiness, /status/startup OK
  • Administration → System Health: all probes green or yellow, none in “unknown” due to missing permissions
Plugins
  • Administration → Plugins: all expected plugins are loaded (incompatible ones are silently skipped – compare with the version list)
  • helpDocs and docx-generator are installed, if the customer is supposed to have them
  • KSeF: resolved
Database and migrations
  • npm run db:migrations – nothing is pending, no unexpected orphan other than Migration20260528120000
  • PostgreSQL: the type of secret_store.valid_from has been verified
  • Certificates: SELECT count(*) FROM secret_store WHERE kind = 'certificate' matches the number of files in the original directory
Crons
  • The cron list matches the pre-upgrade state (active / inactive, timing)
  • No unresolved warnings from parameter normalization in the migration log
  • The cron monitor is being populated; the ES index <index>-cron-runs exists
  • EWS crons: configuration extracted from the [tas-3432] log entries and transferred to MS Graph crons
  • Test manual run of key crons (Run manually)
Logs and monitoring
  • Logs flow into Elasticsearch; no trace of the Arango configuration remains
  • Zabbix scrapes GET /system-health/metrics, LLD discovery works
  • Log filtering by iproc_id / tproc_id / ttask_id / header_id / cron_id works
Frontend
  • Icons are displayed correctly (tas-3710 – a BOM in appStyles.css broke @font-face)
  • Login, dashboard, overviews, case detail, workflow diagram
  • The case diagram (new payload) renders, including archived cases
If the icons are broken, the image comes from a build made before the tas-3710 fix → rebuild it. Emergency workaround: pin "postcss": "8.5.23" in the frontend overrides.
Functional verification
  • Sending e-mail (notification and report)
  • Print / PDF generation (LibreOffice, Puppeteer/Chromium)
  • DMS upload + fulltext + previews
  • AD/LDAP synchronization (AdSyncCron – fixed in 5.19; previously it silently synchronized no one)
  • Templates: saving a task, saving a variable, saving a link in the diagram
  • Login of a user with a bcrypt password → verify the rehash to argon2id
This is a technical smoke test. In-depth verification of templates, calculations, prints and workflow is performed by the consultant according to the consultant part.

Scaling notes

  • Plugin restart (POST /plugins/kill) takes down all backends sharing Redis, including cron workers, in a staggered manner. The deployment must start terminated containers automatically.
  • The plugin volume (TAS_PLUGINS_LOCATION → featureFlag.plugins.destination) must now be writable at runtime and shared by all instances – plugin installation and uninstallation from the GUI write to it.
  • On the Health page, instances are tagged as Web or Cron workers; the instance switcher is in the header.
  • Global infrastructure probes (DB / ES / Redis / Tika / mail / Firebase / crons / business) cache their result in Redis for one sampling window, so that they do not run N times with N instances.

What exactly degrades without Redis (TAS_CACHE_MODE=memory):

Feature

Behavior without Redis

Fleet health registry (HEALTH_FLEET, 48h series)

local instance only

Consecutive cron failure counters, CRONJOB_running stamp, heartbeat alert dedup

does not work across instances

CRON_LAST_RUN-* watermark (PostponedTaskCron, PlanCron)

lost on restart

POST /plugins/kill (pub/sub channel tas.scaling.plugins.kill)

restarting multiple instances does not work

TAS_CACHE_MODE=memory is a fallback mode for single-instance dev, not for production.

CI variables (pipeline only, not runtime): PLUGIN_STORE_URL has been replaced by the pair PLUGIN_STORE_DEV_URL / PLUGIN_STORE_PROD_URL (.github/workflows/image-tag-push.yml) – builds from stable/* publish to the PROD store, all others to DEV. The secret CICD_PLUGIN_STORE_PRIVATE_KEY and the variable PLUGIN_STORE_USERNAME (default tas-ci) remain shared. Publishing is opt-in – until the URL is set, it is a no-op.

Rollback

Migrations delete tables and columns; down() restores only the structure, not the data. A practical rollback therefore looks like this:

  1. Stop 5.19 (web and cron).
  2. Restore the DB from the backup taken in step 2 of the upgrade procedure.
  3. Revert to the 5.17 image and the original env (put TAS_CERTIFICATES_STORAGE_PATH back into the config, and the Arango variables if applicable).
  4. Revert the plugins to their 5.17 versions.
  5. Restore the original Zabbix template.
Restoring the DB is the only reliable path – do not rely on migration down.

Frantisek Brych Updated by Frantisek Brych

Upgrade to TAS 5.19 - Key Changes and Removed Features

Contact

Syca (opens in a new tab)

Powered by HelpDocs (opens in a new tab)