OS: Debian Trixie 13 on a Raspberry PI 5
Grafana: 13.1.1
Loki: 3.7.2 (branch: release-3.7.x, revision: 7486c4a7)
Loki configuration changes for alert state per documentation.
limits_config:
split_queries_by_interval: ‘24h’
max_query_parallelism: 32
Grafana configuration changes for alert state.
[unified_alerting.state_history]
enabled = true
backend = loki
loki_remote_url = http://localhost:3100
[feature_toggles]
#enable = alertingCentralAlertHistory - change to newer method.
alertingCentralAlertHistory = true
I can look at the main history page and not get any loki errors.
I can look at some alert rules with no loki errors.
NO ERRORS:
This one has loki ERRORS:
Loki errors:
| 1784727654113 |
2026-07-22T13:40:54.113Z |
<30>1 2026-07-22T09:40:54.113934-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.055233402Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727654014 |
2026-07-22T13:40:54.014Z |
<30>1 2026-07-22T09:40:54.014076-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.006316047Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727654014 |
2026-07-22T13:40:54.014Z |
<30>1 2026-07-22T09:40:54.014051-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.005943694Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727654014 |
2026-07-22T13:40:54.014Z |
<30>1 2026-07-22T09:40:54.014032-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.004544338Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727654014 |
2026-07-22T13:40:54.014Z |
<30>1 2026-07-22T09:40:54.014013-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.004505153Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727654013 |
2026-07-22T13:40:54.013Z |
<30>1 2026-07-22T09:40:54.013993-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.004408782Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727654013 |
2026-07-22T13:40:54.013Z |
<30>1 2026-07-22T09:40:54.013969-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.003599095Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727654013 |
2026-07-22T13:40:54.013Z |
<30>1 2026-07-22T09:40:54.013886-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:54.003256371Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727644913 |
2026-07-22T13:40:44.913Z |
<30>1 2026-07-22T09:40:44.91362-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.840679261Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814805-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.788957378Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.81479-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.764882582Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814774-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.758363154Z caller=scheduler_processor.go:176 component=querier org_id=fake msg=error notifying scheduler about finished query err=EOF addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814757-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.758140524Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814732-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.757909116Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814711-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.75788606Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
| 1784727644814 |
2026-07-22T13:40:44.814Z |
<30>1 2026-07-22T09:40:44.814611-04:00 raspberry05 loki 230844 - - level=error ts=2026-07-22T13:40:44.757839504Z caller=scheduler_processor.go:111 component=querier msg=error processing requests from scheduler err=rpc error: code = Canceled desc = context canceled addr=127.0.0.1:9096 |
Since these seem to be a Grafana generated issue, is there maybe a configuration that can be added or adjusted?
Since the main Alert History page works, and the History tab also works for another alert rule, this doesn’t appear to be a general alert state history configuration issue.
The issue appears to be specific to this alert rule. I’d first try reducing the History time range (for example, Last 24 hours) and see whether the scheduler_processor.go context canceled/EOF messages still occur. If the behavior changes with a smaller time range, that would suggest the amount of data being queried is a factor.
If the issue is consistently reproducible only for this rule, while other alert history pages continue to work, I’d recommend opening a Grafana or Loki GitHub issue with the rule definition, your Grafana and Loki versions, Loki configuration, steps to reproduce, and the corresponding Loki logs.
How would you reduce the history time range for an alert rule?
What settings and where please?
Do you have any specific information on where to change these settings?
I don’t know what UI you are asking me to use to get there.
What would looking at settings on a phone do to solve my Grafana/Loki issue?
Quick update.
I reduced the Grafana setting loki_max_query_length to 90h from 721h. I have not seen any of the loki errors since. Seems there should be some Loki tuning opportunities that may allow me to increase the Grafana setting back at least some.
Thank you Kimberly. I’m on my way to Owatonna Minnesota now to discuss this further.
Update: Reverted loki_max_query_length back to 721h..
Basically this issue was a result of trying to solve a Grafana issue with
https://community.grafana.com/t/error-saving-alert-annotation-batch/162761/11
I followed a recommendation to send alert state history to Loki.
This is a small home system with a monolithic (single-binary) Grafana Loki deployment.
I had been running well using Loki as a syslog server that maintained 60 days of logs at
about 500MB of data.
Currently I am trying to reduce the amount of state history.
I have started with the following and with adjust to see how much works best.
retention_stream:
- selector: '{from="state-history",group="Prometheus"}'
priority: 1
period: 72h
- selector: '{from="state-history",group="SSH-Connections"}'
priority: 1
period: 72h
- selector: '{from="state-history",group="Rsyslog"}'
priority: 1
period: 72h
Thanks for the update. Since reverting loki_max_query_length to 721h didn’t immediately bring the errors back, it does seem that loki_max_query_length itself wasn’t the root cause.
Limiting retention for the from="state-history" streams is a reasonable next step. If reducing the retained alert state history consistently prevents the scheduler_processor.go context canceled/EOF errors, that would suggest the issue is related to the amount of state-history data being queried.
It would be interesting to see whether the errors remain gone after you’ve run with the shorter retention for a few days.
Just upgraded Loki to latest version.
pi@raspberry05:~ $ loki -version
loki, version 3.7.4 (branch: release-3.7.x, revision: b318f282)
build user: root@9dc424156f20
build date: 2026-07-22T03:46:43Z
go version: go1.26.5
platform: linux/arm64
tags: netgo
Currently waiting for new retention levels to settle.
retention_period: 168h
retention_stream:
- selector: '{from="state-history",group="Prometheus"}'
priority: 1
period: 72h # Prometheus Alerts are wiped after 3 Days
- selector: '{from="state-history",group="SSH-Connections"}'
priority: 1
period: 72h # SSH Connection Alerts are wiped after 3 Days
- selector: '{from="state-history",group="Rsyslog"}'
priority: 1
period: 72h # Rsyslog Alerts are wiped after 3 Days
Current Loki storage is:
pi@raspberry05:~ $ sudo du -h -d 1 /opt/loki
456K /opt/loki/retention
66M /opt/loki/chunks
1.5M /opt/loki/tsdb-shipper-active
632K /opt/loki/wal
968K /opt/loki/tsdb-shipper-cache
4.0K /opt/loki/rules
70M /opt/loki
Down from 500MB.
So far since stabilizing at my current retention levels, I have not seen the Loki errors.
retention_period: 168h and retention_streams set to 72H
caller=scheduler_processor.go:111 component=querier msg="error processing requests from scheduler" err="rpc error: code = Canceled desc = context canceled" addr=127.0.0.1:9096
Current Loki storage usage:
pi@raspberry05:~ $ sudo du -h -d 1 /opt/loki
1.5M /opt/loki/retention
77M /opt/loki/chunks
1.5M /opt/loki/tsdb-shipper-active
536K /opt/loki/wal
1.1M /opt/loki/tsdb-shipper-cache
4.0K /opt/loki/rules
82M /opt/loki
This is my current Alert State History levels.
I am going to attempt to set retention levels as follows.
retention_period: 720h
retention_stream:
- selector: '{from="state-history",group="Prometheus"}'
priority: 1
period: 168h # Prometheus Alerts are wiped after 7 Days
- selector: '{from="state-history",group="SSH-Connections"}'
priority: 1
period: 168h # SSH Connection Alerts are wiped after 7 Days
- selector: '{from="state-history",group="Rsyslog"}'
priority: 1
period: 168h # Rsyslog Alerts are wiped after 7 Days
```
So I am inclined to believe how Grafana queries Loki when you look at the history page
of an alert is what is causing the error. I don’t see the error when looking at the main history
page as it defaults to just 1 hour for a time period.
A few hours later.
When looking at alert history for Rsyslog alerts.
Only a small increase in records for history of this alert.
Updated to latest version.
Upgraded Loki:
pi@raspberry05:~ $ loki -version
loki, version 3.7.6 (branch: release-3.7.x, revision: 5003600d)
build user: root@31d37f362fe2
build date: 2026-08-05T15:36:54Z
go version: go1.26.5
platform: linux/arm64
tags: netgo
Current Loki storage:
pi@raspberry05:~ $ sudo du -h -d 1 /opt/loki
456K /opt/loki/retention
118M /opt/loki/chunks
1.5M /opt/loki/tsdb-shipper-active
500K /opt/loki/wal
1.2M /opt/loki/tsdb-shipper-cache
4.0K /opt/loki/rules
122M /opt/loki
Still waiting on new streams retention to fill out. (168h)
Current Alert State History levels:
So the problem with the Loki errors I have been seeing when looking at the history of
a alert appear the be caused by how Grafana queries the data.
If I look at the main history page for all alerts I do not get any Loki errors even when selecting
a 7 day time period. (My current streams retention period.) It just displays a message box
telling you to reduce your request.

While if I look at the history of my Rsyslog alert I get Loki errors.
This is my current Alert state history.
I guess this is time to post this on Grafana github to see if it can be corrected.