Senior Site Reliability Engineer - AEM - Content Delivery Network
Astra North Infoteck Inc. · Toronto, Canada
Senior Site Reliability Engineer - AEM - Content Delivery Network
Toronto- 4 Days WFO
ABOUT THE ROLE
1. Application Support & Incident Management
• Own end-to-end monitoring of the controlled surface: CDN
and edge configuration, DNS, certificates, cache and invalidation health, and
every third-party integration on the page – Search, Consent Management,
Analytics, Personalization and AI services and many more to come.
• Build and run synthetic monitoring from outside the bank
network, per template, per language, because internal-only monitoring cannot
see the CDN, DNS and certificate failures this architecture is most exposed to.
• Run smoke testing of dependent interfaces on every change
and maintain the automation packs that do it.
• Participate in the shared on-call rotation as the
platform’s subject-matter escalation, and lead incident management for
customer-facing events.
• Own the vendor's escalation path: severity mapping between
vendor and internal incident scales, named contacts, evidence capture, and
holding the vendor to its commitment during an event.
• Handle a class of incident that does not exist on
traditional platforms – content published but not visible, invalidation
failure, and authoring-source outages – and make those diagnosable by the
service desk rather than by you.
2. Change and Release Reliability
• Design and operate change management for the platform
where the Git repository is production: reconcile a merge-to-main deployment
model with change control, so that every production change carries an approved
record with stalling delivery.
• Own the release pipeline as a production control – branch
protection, required checks, lint, performance, and secret-screening gates –
and the evidence that they are enforced.
• Own rollback: revert, republish and purge, rehearsed end
to end with a measured recovery time and a named authority who can call it
without convening a meeting.
• Treat content publishing as a routine process: hundreds of
production changes made by content authors, needing approval evidence,
attribution and retention trail.
• Represent the platform at change advisory board, and own
the freeze calendar interaction and release notes.
3. Business Continuity and Resilience
• Own the recovery obligation. The vendor operates delivery
resiliently, but customers restore their own content from source version
history rather than vendor backups – so the content source, the Git repository
and the CDN configuration are the recovery surface, each needing a tested
restore.
• Hold CDN and edge configuration as code so that a lost or
corrupted property is a redeploy rather than an outage with no runbook.
• Define RTO and RPO with the business against the
application criticality tier, document the DR exercise plan, and execute the
testing – including failure modes you can actually cause: certificate expiry,
invalidation failure, WAF misconfiguration, content source unavailability, and
repository compromise.
• Maintain the operational resilience evidence a regulator
expects for a material third-party technology arrangement, and keep the
platform exit and portability plan current.
4. Reliability & Performance Engineering
• Set and defend service level objectives for both
availability and page performance. Define Core Web Vitals thresholds per
template, run them on an error budget, and report against them.
• Build the observability practice from the telemetry that
exists; real user monitoring on the production domains, CDN access logs
streamed to enterprise SIEM as the log source of record, and external
synthetics. There is no origin server log – designing around that constraint is
part of the job.
• Own third-party scripts and tag governance as a
reliability control. Tags are the dominant cause of performance regressions and
are added by teams outside engineering change control; you will define the
approval route, measure each tag’s cost and enforce the budget.
• Own capacity and cost where they still exist: CDN egress,
asset storage and processing, media delivery and any hosted APIs behind the
page. Capacity planning here is a financial operations discipline, not a
server-sized one.
• Publish the reliability and performance reporting that the
business, risk and technology leadership use.
5. Compliance and Control Evidence
• Evidence controls on a platform the organization does not
operate – which is harder than evidencing your own, and is where a meaningful
share of the role’s effort sits.
• Own log ingestion into SIEM with the agreed retention,
access recertification across the repository, admin console, content source and
CDN, and the audit evidence pack.
• Support privacy, operational risks, control assessment and
third-party risk processes with operational evidence and maintain alignment to
regulatory expectations for technology, cyber and third-party risks.
• Keep the configuration management database, support model
and assignment groups accurate as the platform estate grows.
WHAT WILL YOU DO?
This is a build-then-run role. Roughly half of the first
year is establishing a reliability practice that does not exist yet.
• The observability stack: Real User Monitoring, Core Web
Vitals dashboards and alerting, CDN log ingestion, and external synthetics.
• CDN and Edge configuration as Code, with a tested restore.
• The operations runbook, incident playbook, operational
level agreement and vendor escalation matrix.
• The change model that reconciles Git-based deployments and
continuous content publishing with change controls.
• The first disaster recovery exercise and the first
rehearsed, measured rollback.
• Service level objectives agreed with the business, and the
reporting that holds the platform to them.
WHAT DO YOU NEED TO SUCCEED?
Must have
• Substantial hands-on experience operating a high-traffic
public website behind an enterprise content delivery network. Depth in CDN
configuration – origin and cache behaviour, invalidation, edge logic, TLS and
DNS – is the single most important qualification. Akamai and Cloudflare
experience is an advantage.
• Practical web application firewall experience, including
tuning false positives against production-like traffic before enforcement, and
bot management that protects the site without blocking the crawlers you need.
• A real observability practice: defining service level
objectives and error budgets, and building monitoring from log, real-user and
synthetic sources rather than from an agent on a server.
• Web performance engineering – Core Web Vitals, load and
rendering behaviour, and the ability to read a waterfall and attribute a
regression to a specific script.
• Comfortable with front-end technology: This platform ships
JavaScript and CSS to the browser with no server tier; you cannot reason about
its reliability without reading and understanding it.
• Git-based release engineering and CI/CD as a production
control, including infrastructure and configuration as code.
• Incident command on customer-facing services, and the
discipline to produce evidence during an event, not after it.
• Working effectively in a regulated environment – change
control, audit evidence, access management and third-party risk – without
treating it as an obstacle.
Nice-to-have
• Experience operating a vendor-run or SaaS-delivered
platform, where reliability means instrumenting, escalating and holding a
supplier accountable rather than fixing the tier yourself.
• Adobe Experience Manager exposure, particularly Edge
Delivery Services and Assets as a Cloud Service.
• Financial services or another regulated sector.
• Bilingual delivery – operating a site that must meet the
same standard in English and French.
• Accessibility and Search Engine Optimization literacy
sufficient to recognize when a reliability decision creates a compliance or
discoverability problem.
• Automation in Python, Java or JavaScript, and a preference
for encoding a runbook rather th