Site Reliability Engineer (SRE)

Job Description
SS&C Technologies is a leading financial services and healthcare technology company headquartered in Windsor, Connecticut. With 27,000+ employees in 35 countries, SS&C serves 20,000+ organizations worldwide, ranging from large enterprises to mid-market firms, providing cutting-edge technology and expertise.
About the Team
The Intralinks SRE Team is responsible for ensuring the availability, scalability, and reliability of Intralinks’ production application platform. The team continuously improves monitoring and operational efficiency to deliver uninterrupted services to clients.
Job Overview
As an SRE, you will play a crucial role in analyzing root causes of incidents, managing operational tasks, and automating monitoring and alerting systems to minimize Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR).
Requirements
- Respond to and resolve escalated incidents related to customer issues or monitoring alerts.
- Conduct in-depth root cause analysis of incidents and implement long-term solutions.
- Collaborate with R&D and architecture teams to resolve production defects and inefficiencies.
- Build diagnostic tools and automation to improve MTTA, MTTD, and MTTR.
- Develop and integrate monitoring tools to ensure system health and API performance.
- Validate and verify software deliverables for production readiness.
- Assess risks and mitigate potential impacts of changes in the production environment.
Minimum Qualifications & Experience
- Strong experience in Unix/Linux environments and troubleshooting skills.
- Expertise in Java-based enterprise applications and performance tuning.
- Hands-on experience with Kubernetes for managing microservices and containerized applications.
- Working knowledge of AWS services, including CloudWatch, EKS, EFS, S3, and RedShift.
- Experience with AWS tools for troubleshooting application constraints, connectivity, monitoring, and alerting.
- Proficiency in one or more programming languages (Java, Python, etc.) for automation and system tuning.
- Scripting experience (Shell, Python, etc.) for automating tasks.
- Strong understanding of networking concepts and application/transport protocols (HTTP, JMS, TCP, UDP).
- Experience with monitoring tools such as Splunk, Dynatrace, Zabbix, and Prometheus.
- Database expertise in Oracle, PostgreSQL, or MongoDB.
- Experience with messaging systems (RabbitMQ, AMQ, Interconnect).
- Proficiency in CI/CD tools such as Jenkins and Git.
Location & Salary
- Location: Hyderabad, India
- Salary: Competitive, based on experience
Why Join SS&C Technologies?
- Work in a global, fast-paced environment at the forefront of financial and healthcare technology.
- Opportunities to innovate and implement SRE best practices across a mission-critical platform.
- Career development and continuous learning opportunities with cutting-edge technologies.
- Competitive salary and comprehensive benefits package.
- A diverse and inclusive work environment that fosters innovation and growth.
SS&C Technologies is an Equal Employment Opportunity employer and does not discriminate based on race, gender, age, disability, or any other protected category.
Interested? Apply today and be part of a world-class SRE team ensuring high availability and reliability of mission-critical systems!