Position Title: Hybrid Cloud & Identity Engineer
Reports To: Senior Director, Hybrid Cloud, Identity and Infrastructure
Department: Information Technology
Leading Retail Organization
Summary
The Hybrid Cloud & Identity Engineer supports, maintains, and improves the enterprise hybrid cloud, identity, server, storage, and virtualization environment. This role blends hands-on operational support with modern cloud engineering practices across Microsoft Azure, Entra ID, Active Directory, Windows Server, Red Hat Enterprise Linux, SAN storage, hypervisors, backup/recovery, monitoring, and automation.
The engineer is expected to work independently, execute approved changes with strong validation, troubleshoot complex incidents across infrastructure layers, produce clear documentation, and help advance secure, resilient, cost-conscious hybrid cloud operations.
Key Responsibilities
Hybrid Cloud and Azure Administration / Architecture:
- Support and administer Azure infrastructure resources, including subscriptions, resource groups, virtual machines, storage accounts, virtual networks, NSGs, routing dependencies, Azure Monitor, backup, recovery, and Azure Arc-enabled servers.
- Contribute to Azure landing zone, governance, tagging, RBAC, resource lifecycle, resiliency, and cost-control standards in partnership with infrastructure, security, network, finance, and application teams.
- Assist with cloud migration and modernization work, including workload assessment, resource planning, validation testing, cutover support, and post-migration operational readiness.
- Apply secure-by-design and least-privilege practices when supporting Azure and identity-related resources.
Identity and Microsoft Platform Engineering:
- Administer and support Active Directory, Entra ID, Group Policy, DNS, DHCP, LDAP, PKI-related dependencies, hybrid identity synchronization, authentication, authorization, SSO, MFA, Conditional Access, and Privileged Identity Management.
- Support selected Microsoft 365 and E5 ecosystem services, including Exchange Online, SharePoint Online, Teams, OneDrive, Intune, Defender, Identity Protection, and related governance controls within assigned scope.
- Support Azure Virtual Desktop environments, including host pool health, session host troubleshooting, FSLogix dependencies, capacity validation, monitoring, and patching coordination.
- Maintain directory hygiene, object lifecycle processes, access troubleshooting workflows, and identity-related operational documentation.
Red Hat Linux and Server Lifecycle Management:
- Administer Red Hat Enterprise Linux servers, including provisioning support, patching, package management, systemd services, user/group management, permissions, filesystems, LVM, networking, SSH, logs, and core troubleshooting.
- Support RHEL lifecycle planning, version upgrades, EOL/EOS remediation, vulnerability remediation, configuration standards, and compliance evidence collection.
- Use Red Hat tools and ecosystem capabilities where applicable, such as Red Hat Satellite, Insights, repositories, subscription management, SELinux awareness, and approved shell scripting practices.
- Build, patch, upgrade, monitor, and maintain Windows and Linux servers according to enterprise standards, approved changes, and security requirements.
SAN Storage, Hypervisor, and Data Center Infrastructure:
- Support enterprise SAN storage concepts and operational workflows, including capacity awareness, LUN presentation, zoning coordination, multipathing dependencies, replication awareness, and escalation with storage vendors or senior engineers.
- Support virtualization and private cloud platforms such as VMware vSphere/ESXi, Nutanix AHV/HCI, Hyper-V, or comparable hypervisors, including VM provisioning, resource validation, cluster health, host maintenance, and incident troubleshooting.
- Coordinate with hardware, facilities, network, backup, and application teams for physical and virtual server lifecycle work, data center activities, maintenance windows, and dependency validation.
- Participate in storage, hypervisor, and backup/recovery validation for production, development, test, and disaster recovery environments.
Security, Compliance, Change, and Incident Response:
- Execute scheduled vulnerability remediation, patching, operating system upgrades, monitoring updates, and approved infrastructure changes with appropriate validation and rollback awareness.
- Respond to alerts and incidents, isolate issues across identity, compute, network, storage, virtualization, backup, and application dependency layers, and escalate business-impacting risks promptly.
- Provide support for PCI, SOX, audit, attack-and-penetration remediation, security control implementation, evidence collection, and change management processes.
- Participate in post-incident reviews and contribute corrective actions, monitoring improvements, and runbook updates.
Automation, Documentation, and Continuous Improvement:
- Develop and maintain repeatable scripts, runbooks, standards, and procedures using PowerShell, Azure PowerShell, Azure CLI, shell scripting, Terraform, Bicep, ARM, Ansible, Puppet, Git, or comparable approved tooling.
- Improve documentation quality across SOPs, knowledge base articles, diagrams, inventories, configuration records, handoff notes, and operational procedures.
- Identify opportunities to reduce manual work, improve build consistency, strengthen operational readiness, enhance monitoring, and improve self-service request patterns.
- Collaborate effectively with cybersecurity, networking, database, application, desktop support, procurement, finance, vendors, and infrastructure teams.
Required Qualifications
- 5+ years of experience supporting enterprise infrastructure, identity, cloud, Windows Server, Linux, virtualization, storage, or hybrid operations environments.
- 3+ years of hands-on Azure administration experience with resources such as virtual machines, virtual networks, subscriptions, resource groups, storage, monitoring, backup, RBAC, policies, and Azure Arc.
- Working knowledge of Azure architecture patterns, including landing zones, network segmentation, identity integration, governance, resiliency, lifecycle management, and cost-aware design.
- Strong experience administering Active Directory, Entra ID, Group Policy, DNS, DHCP, LDAP, authentication, authorization, MFA, Conditional Access, and privileged access concepts.
- Hands-on Red Hat Enterprise Linux administration experience, including patching, package management, shell access, system services, storage/filesystem basics, logs, networking, permissions, and troubleshooting.
- Experience supporting hypervisor platforms such as VMware ESXi/vSphere, Nutanix AHV/HCI, Hyper-V, or equivalent virtualization technologies.
- Working knowledge of SAN storage concepts, including LUNs, zoning, host mappings, multipathing, replication dependencies, capacity management, and storage-related incident triage.
- Experience with vulnerability remediation, patch management, monitoring, backup validation, incident troubleshooting, documentation, and formal change control.
- Ability to work independently during maintenance windows or low-staffed coverage periods while making appropriate escalation decisions.
- Strong written and verbal communication skills with the ability to translate technical issues into clear business and operational language.
Preferred Qualifications
- Microsoft Azure Administrator Associate, Azure Solutions Architect Expert, Microsoft Identity, RHCSA, RHCE, VMware, Nutanix, ITIL, FinOps, or related infrastructure certifications.
- Experience with Red Hat Satellite, Red Hat Insights, SELinux, Ansible, Puppet, Terraform, Bicep, Git-based workflows, REST APIs, or Microsoft Graph.
- Experience with Azure Virtual Desktop, FSLogix, Intune, MECM, Defender, Exchange Online, SharePoint Online, Teams, OneDrive, and Microsoft 365 E5 security capabilities.
- Experience with Dell/EMC, IBM, HPE, Brocade, Rubrik, Nutanix, VMware/Broadcom, or comparable enterprise infrastructure platforms.
- Experience supporting cloud migrations, data center modernization, storage refreshes, hypervisor transitions, OS lifecycle programs, disaster recovery, or large infrastructure transformations.
- Familiarity with observability and monitoring platforms such as Datadog, Splunk, Microsoft Sentinel, Dynatrace, Prometheus, Grafana, Zabbix, or comparable tools.
Success Measures
- Reliable execution of approved maintenance windows, deployments, patching, and validation tasks.
- Measurable progress on server, identity, Azure, Linux, storage, hypervisor, lifecycle, and compliance hygiene.
- Reduced incident recurrence through better documentation, monitoring quality, root-cause follow-up, and operational standards.
- High-quality handoffs across teams, vendors, and coverage windows.
- Demonstrated ability to work independently while escalating appropriately when business impact, security risk, or architectural decision points are present.