Event Processing Monitoring (procevent)
procevent is the Altcraft process that handles events (clicks, opens, deliveries, hooks, and others) arriving via the ak_proc_event RabbitMQ queue. The process supports Prometheus metrics, batch processing mode, and batch event insertion into ClickHouse — tools for diagnosing and optimizing performance under high load.
Metrics setup
To enable procevent metrics, add the process to the PROMETHEUS_METRICS.PROCESSES list and specify the host and port of the metrics HTTP server:
{
"PROMETHEUS_METRICS": {
"ENABLE": true,
"PROCESSES": ["procevent"]
},
"PROC_EVENT_METRIC_HOST": "0.0.0.0",
"PROC_EVENT_METRIC_PORT": 8916
}
| Parameter | Type | Default | Description |
|---|---|---|---|
PROC_EVENT_METRIC_HOST | string | localhost | Host of the Prometheus metrics HTTP server |
PROC_EVENT_METRIC_PORT | int | 8916 | Port of the Prometheus metrics HTTP server |
Add a scrape job to the Prometheus configuration:
scrape_configs:
- job_name: 'procevent'
metrics_path: /metrics
static_configs:
- targets:
- '<procevent-server>:8916'
All metrics use the altcraft_procevent_ prefix.
Available metrics
Duration metrics
| Metric | Type | Labels | What it measures |
|---|---|---|---|
event_processing_duration_seconds | Histogram | group | Total time to process an event, from RabbitMQ delivery to the batch writer write |
find_lead_duration_seconds | Histogram | — | MongoDB lookup on lead cache miss |
unique_action_duration_seconds | Histogram | action_type | Unique action insert (delivery or popup dedup) |
hook_check_duration_seconds | Histogram | — | Hook existence check in the action cache |
hook_send_duration_seconds | Histogram | hook_group | Sending a hook action to actionSender (blocks if the queue is full) |
mongo_update_send_duration_seconds | Histogram | — | Sending a profile update to the mgoUpdater channel (blocks if the channel is full) |
ssdb_hb_send_duration_seconds | Histogram | — | Sending a heartbeat to the ssdbHBSupp channel (blocks if the channel is full) |
trigger_write_duration_seconds | Histogram | trigger_type | Writing a trigger to trigCache |
ch_batch_insert_duration_seconds | Histogram | batcher | Inserting a batch into ClickHouse (prepare + exec per event + commit) |
ch_insert_send_duration_seconds | Histogram | batcher | Sending an event to the batch writer insertChan (blocks if the channel is full) |
ch_batch_size | Histogram | batcher | Batch size at flush time |
ack_call_duration_seconds | Histogram | — | Time spent in the consumer.Ack() call (RabbitMQ acknowledgment) |
cache_refresh_duration_seconds | Histogram | cache | Time spent refreshing cache data from MongoDB |
Counters
| Metric | Type | Labels | What it counts |
|---|---|---|---|
events_processed_total | Counter | group, result | Processed events by group and result (ok / error / retry) |
lead_cache_lookup_total | Counter | result | Lead cache lookups (hit / miss) |
gcg_cache_lookup_total | Counter | result | gcg cache lookups (hit / miss) |
hook_check_total | Counter | result, event_type | Hook checks (hit / miss) by event type |
custom_channel_lookup_total | Counter | result | Custom channel cache lookups (hit / miss) |
hook_send_total | Counter | hook_group | Hook actions sent by group |
trigger_writes_total | Counter | trigger_type | Triggers written by type |
unique_action_duplicates_total | Counter | action_type | Duplicates on unique action insert (expected dedup, not an error) |
rmq_events_total | Counter | queue | Events consumed from RabbitMQ |
rmq_nack_total | Counter | queue | Events returned to RabbitMQ (nack, will be retried) |
ch_tx_retries_total | Counter | batcher | ClickHouse transaction retries (reconnection attempts) |
ch_event_errors_total | Counter | batcher | Per-event exec failures during ClickHouse batch insert |
ch_commit_errors_total | Counter | batcher | ClickHouse transaction commit failures |
Batch processing mode
Batch mode collects messages from the queue into batches of up to 4000 events and processes them as a group.
Disabled by default: false
| Parameter | Type | Default | Description |
|---|---|---|---|
PROC_EVENT_ENABLE_RMQ_BATCH_MODE | bool | false | Enables batch processing mode for events |
Enabling changes the processing pipeline:
- Default (stable mode): each event is processed individually: queue → 1 event → processing → write to the insertChan → asynchronous batch insertion into ClickHouse.
- Batch mode: messages are collected into a batch: queue → batch → batch find-lead (1 MongoDB query) → process each event → synchronous batch insert into ClickHouse (1 transaction) → split ACKs.
In batch mode the PROC_EVENT_WORKER_COUNT parameter is ignored — the worker count is forced to 1. Parallelism is achieved only via PROC_EVENT_CONSUME (the number of independent consumers).
Batch mode metrics
| Metric | Type | Labels | What it measures |
|---|---|---|---|
batch_events_total | Counter | result | Events processed in batch mode by result |
batch_ch_insert_seconds | Histogram | — | Time to insert a batch into ClickHouse |
batch_mongo_find_seconds | Histogram | — | Time of the batch find-lead in MongoDB |
batch_rmq_collect_seconds | Histogram | — | Time to collect the batch from the queue |
batch_ack_seconds | Histogram | — | Time of the split ACKs |
batch_size | Histogram | — | Batch size |
batch_ack_groups_total | Counter | type | Number of ACK/NACK groups |
batch_mongo_groups_total | Histogram | — | Number of batch find groups |
batch_mongo_queries_total | Histogram | — | Number of batch find queries |
resolve_lead_total | Counter | source | Lead source (cache, MongoDB, etc.) |
Native ClickHouse batch
The native batch switches the batch event insertion into ClickHouse from the SQL protocol to the native binary protocol.
| Parameter | Type | Default | Description |
|---|---|---|---|
PROC_EVENT_ENABLE_NATIVE_BATCH | bool | false | Switches batch event insertion into ClickHouse to the native binary protocol |
Mechanism difference:
- SQL protocol (default): transaction begin → prepare → exec loop per event → commit.
- Native binary protocol: PrepareBatch → Append → Send (ClickHouse binary protocol) — no per-event exec commands.
Recommended configuration for high load
Based on load testing:
{
"PROC_EVENT_CONSUME": 4,
"PROC_EVENT_PREFETCH_SIZE": 64000,
"PROC_EVENT_WORKER_COUNT": 8,
"CLICKHOUSE_BATCH_WORKERS_SIZE": 4,
"RMQ_CLICKHOUSE_BLOCK_SIZE": 8000,
"SEGMENT_ACTION_GROUP_OPTIMIZATION": true,
"PROC_PIPER_MESSAGES_PREFETCH_SIZE": 32000,
"PIPE_ROUTER_WORKER_SIZE": 32,
"PIPE_WORKER_SIZE": 64,
"FIREBASE_PUSH_CONSUMER_PREFETCH_SIZE": 16000
}
| Parameter | Type | Default | Description |
|---|---|---|---|
PROC_EVENT_CONSUME | int | 1 | Number of independent queue consumers |
PROC_EVENT_PREFETCH_SIZE | int | -1 | Consumer prefetch window size |
PROC_EVENT_WORKER_COUNT | int | — | Number of processing workers (ignored in batch mode) |
CLICKHOUSE_BATCH_WORKERS_SIZE | int | 8 | Number of ClickHouse batch writer workers |
RMQ_CLICKHOUSE_BLOCK_SIZE | int | — | RabbitMQ → ClickHouse block size |
SEGMENT_ACTION_GROUP_OPTIMIZATION | bool | false | Segment action grouping optimization |
PROC_PIPER_MESSAGES_PREFETCH_SIZE | int | 500 | Piper messages queue prefetch window |
PIPE_ROUTER_WORKER_SIZE | int | 1 | Number of pipe router workers |
PIPE_WORKER_SIZE | int | 1 | Number of pipe processing workers |
FIREBASE_PUSH_CONSUMER_PREFETCH_SIZE | int | -1 | Firebase Push consumer prefetch window |
Problem diagnostics
| Problem | Indicators |
|---|---|
| Channel blocking | hook_send_duration_seconds p95 > 1 s, mongo_update_send_duration_seconds p95 > 1 s, ssdb_hb_send_duration_seconds p95 > 1 s |
| MongoDB problems | high find_lead_duration_seconds p99, low hit rate in lead_cache_lookup_total |
| ClickHouse problems | high ch_batch_insert_duration_seconds, growing ch_tx_retries_total, growing ch_commit_errors_total |
| Batch writer lag | ch_insert_send_duration_seconds > 0, small ch_batch_size |
| Throughput | rate of events_processed_total broken down by group |
Grafana dashboard
A ready-made Grafana dashboard for procevent metrics is available: procevent-dashboard-en.json.
To import it:
1. Open Grafana → Dashboards → New → Import.
2. Upload the procevent-dashboard-en.json file.
3. Select your Prometheus data source in the Prometheus field.
4. Click Import.