|
2026-09-23
§
|
| 09:59 |
<ayounsi@cumin1004> |
START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004 |
[production] |
| 09:58 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet |
[production] |
| 09:58 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet |
[production] |
| 09:58 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet |
[production] |
| 09:57 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:55 |
<brouberol@cumin1004> |
DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0) |
[production] |
| 09:54 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:54 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:53 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:52 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:51 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:51 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet |
[production] |
| 09:51 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet |
[production] |
| 09:51 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet |
[production] |
| 09:51 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet |
[production] |
| 09:50 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:44 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet |
[production] |
| 09:43 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet |
[production] |
| 09:43 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{dse-k8s-wdqs100[1-3].eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) |
[production] |
| 09:38 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 09:34 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:45 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:44 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:44 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:41 |
<brouberol@cumin1004> |
DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:27 |
<brouberol@cumin1004> |
END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99) |
[production] |
| 08:27 |
<brouberol@cumin1004> |
START - Cookbook sre.ceph.remove-osd |
[production] |
| 08:25 |
<kevinbazira@deploy1003> |
helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . |
[production] |
| 08:24 |
<kevinbazira@deploy1003> |
helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . |
[production] |
| 08:13 |
<tappof@deploy1003> |
Finished scap sync-world: T432444 - Provision kafka-logging100[6-8] (duration: 12m 52s) |
[production] |
| 08:05 |
<moritzm> |
installing grub2 bugfix updates on Bookworm hosts |
[production] |
| 08:04 |
<tappof@deploy1003> |
Started scap sync-world: T432444 - Provision kafka-logging100[6-8] |
[production] |
| 08:00 |
<tappof@deploy1003> |
helmfile [codfw] DONE helmfile.d/admin 'sync'. |
[production] |
| 07:59 |
<tappof@deploy1003> |
helmfile [codfw] START helmfile.d/admin 'sync'. |
[production] |
| 07:59 |
<moritzm> |
installing giflib security updates |
[production] |
| 07:58 |
<tappof@deploy1003> |
helmfile [eqiad] DONE helmfile.d/admin 'sync'. |
[production] |
| 07:58 |
<tappof@deploy1003> |
helmfile [eqiad] START helmfile.d/admin 'sync'. |
[production] |
| 07:29 |
<moritzm> |
installing python-idna security updates |
[production] |
| 02:08 |
<mwpresync@deploy1003> |
Finished scap build-images: Publishing wmf/next image (duration: 07m 39s) |
[production] |
| 02:00 |
<mwpresync@deploy1003> |
Started scap build-images: Publishing wmf/next image |
[production] |
| 00:50 |
<ryankemper@cumin2003> |
END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance |
[production] |
| 00:46 |
<ryankemper> |
[WDQS] T435443 Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture |
[production] |
| 00:45 |
<ryankemper@cumin2003> |
START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance |
[production] |