Jump to content

Wikimedia Release Engineering Team/Pretrain/Progress reports/2026-06-05

From mediawiki.org

Report on activities in the Pretrain project for the week ending 2026-06-05.

[WE6.7.2] Routing test

[edit]

If we create the Pretrain deployment environment using existing production MediaWiki container images and begin routing testwiki traffic to it, we will develop confidence that our routing design is sufficiently complete to support the full Pretrain MVP (e.g., all testwiki traffic is reliably served by the Pretrain environment and container).

Progress update
  • Stage: Engineering/Development
  • [Limited progress due to PTO and travel this week.]
  • This week focused on refining implementation details, particularly for routing of testwiki-bound API traffic from internal clients that do not use the standard service mesh (see overview in T427666). While further study is needed, these workloads may influence our decision on where / how we will divert traffic (i.e., whether we need to revisit adopting k8s Ingress more widely to centralize diversion).
Any new metrics related to the hypothesis
  • None
Any emerging blockers or risks
  • Not yet
Any unresolved dependencies - do you depend on another team that hasn’t already given you what you need? Are you on the hook to give another team something you aren’t able to give right now?
  • No
Have there been any new lessons from the hypothesis?
  • Not yet
Have there been any changes to the hypothesis scope or timeline?
  • No

[WE6.7.3] Automated supervision

[edit]

If we implement an initial set of automated supervision strategies and enable automated deployment to the Pretrain environment for the routing test, we will develop confidence that automated supervision is ready to serve as our primary risk mitigation for the full Pretrain MVP (e.g., false positives are understood, and we have identified a path to mitigate them).

Progress update
  • Stage: Design
  • Tyler proposed leveraging Alertmanager configuration data (example). Bryan wonders if we can go even further and actually use Alertmanager to define the monitoring thresholds and send webhook information to SpiderPig when an alert fires. Discussion with Scott and Ahmon identified some possible pros and cons with both ideas. We may have an easier time deciding on possibilities here when we have a more concrete list of golden signals (SLIs) that we want to monitor.
  • Ahmon will be out of office next week.
Any new metrics related to the hypothesis
  • None
Any emerging blockers or risks
  • Organizational turmoil such as unexpected staff departures, team reorganizations, and program changes create distractions that detract from productive work. Mental health days, monitoring community discussions, and the cognitive load of preparing for, participating in, and debriefing after group meetings exploring the ongoing situations are time away from the hypothesis work. This risk is not local to this hypothesis or team, but is instead endemic for the organization.
Any unresolved dependencies - do you depend on another team that hasn’t already given you what you need? Are you on the hook to give another team something you aren’t able to give right now?
  • No
Have there been any new lessons from the hypothesis?
  • Not yet
Have there been any changes to the hypothesis scope or timeline?
  • No