1. Job Description
1.1 System Operations and Monitoring
- Monitor the operational status of systems, applications, servers, databases, and related services.
- Perform periodic system health checks (Daily/Weekly/Monthly Health Checks) to ensure stable and continuous system operations.
- Monitor CPU, Memory, Disk, Services, Applications, Databases, and other key operational metrics.
- Check the status of scheduled jobs, batch jobs, services, and automated processes.
- Receive and handle alerts generated by system monitoring tools.
1.2 Incident Handling and System Support
- Receive, classify, and resolve system-related Incidents/Problems.
- Analyze logs, error messages, and system status to identify the root causes of incidents.
- Perform troubleshooting and restore systems within the required timeframe.
- Coordinate with Application, Infrastructure, Network, Database, Security teams, and Vendors to resolve complex incidents.
- Perform Root Cause Analysis (RCA) for critical or recurring incidents.
1.3 Deployment and Change Management
- Perform and coordinate the deployment of releases, patches, configurations, and system changes across Test/UAT/Staging/Production environments.
- Conduct pre- and post-deployment system checks to ensure that all services are operating normally.
- Manage and execute Change Requests in accordance with established procedures.
- Coordinate rollback activities when issues occur after deployment.
- Monitor and assess the impact of system changes on system performance and business operations.
1.4 Data, Batch & Integration Management
- Monitor and ensure that batch jobs, data processing tasks, and scheduled tasks are executed on time and produce the expected results.
- Check data synchronization between internal systems and third-party systems.
- Detect, investigate, and resolve issues related to data, interfaces, and integrations.
- Perform data validation and data reconciliation to ensure data completeness and accuracy.
- Coordinate with Database/Application teams to resolve data-related issues arising during system operations.
1.5 Backup, Recovery & Business Continuity
- Monitor and verify the backup status of systems, databases, and critical data.
- Perform or coordinate testing of data restore and recovery capabilities.
- Participate in developing and implementing Disaster Recovery (DR), Business Continuity, and System Recovery plans.
- Ensure that systems can be restored in the event of incidents or failures.
- Periodically test backup/recovery plans and document the results.
1.6 Security & System Operations Management
- Manage user accounts, access permissions, and system access in accordance with company policies.
- Monitor unusual activities, security alerts, and system security-related issues.
- Ensure that system operations comply with security policies and internal regulations.
- Coordinate with Security/Infrastructure teams in handling security incidents and vulnerabilities.
- Manage operational documentation, SOPs, Runbooks, Operation Checklists, and system configuration information.
1.7 Reporting, Improvement & Operations Optimization
- Monitor and report operational metrics such as system availability, incidents, performance, batch status, and system health.
- Analyze incidents and operational issues to propose preventive measures and improve system stability.
- Automate operational tasks, monitoring, reporting, and repetitive processes.
- Improve operational, troubleshooting, and incident management processes to reduce incident resolution time.
- Research and recommend new tools and solutions to enhance System Availability, Performance, Reliability, and Operational Efficiency.
2. Requirements
Experience & System Operations Knowledge
- College/University degree in Information Technology, Computer Science, Information Systems, or related fields.
- Experience in System Operations, Application Support, IT Operations, or System Administration.
- Knowledge of operational processes, monitoring, incident management, problem management, and change management.
- Understanding of the system operations lifecycle, from deployment and monitoring to incident handling and maintenance.
Systems & Infrastructure Knowledge
- Knowledge of Windows/Linux Servers, Networks, Databases, and Application Servers.
- Ability to check system resources such as CPU, Memory, Disk, Services, and Network Connectivity.
- Experience using Monitoring, Log Management, and System Management tools.
- Ability to read and analyze system/application logs to identify the causes of incidents.
Troubleshooting & Problem-Solving Skills
- Ability to receive, classify, prioritize, and handle Incidents/Service Requests.
- Strong analytical and troubleshooting skills, with the ability to identify the root causes of incidents.
- Ability to handle incidents in Production environments and restore systems within the required timeframe.
- Ability to perform RCA and recommend preventive measures to avoid recurring incidents.
Database, Data & System Integration
- Basic knowledge of SQL Server, PostgreSQL, MySQL, or equivalent database management systems.
- Ability to use SQL for data checking, reconciliation, and queries required for incident resolution.
- Understanding of APIs, Web Services, Data Integration, and System Interfaces.
- Ability to investigate and resolve data synchronization issues between systems.
Deployment, Backup & Recovery
- Experience with or understanding of Deployment/Release Management processes across Test/UAT/Staging/Production environments.
- Knowledge of Backup, Restore, Recovery, and Disaster Recovery.
- Understanding of Change Management processes and risk controls during system deployments.
- Ability to perform pre- and post-deployment system checks.
Security & Operational Compliance
- Basic knowledge of System Security, User Access, Authentication, and Authorization.
- Understanding of access management principles and data protection in enterprise environments.
- Strong awareness of and compliance with SOPs, Security Policies, IT Policies, and operational procedures.
- Ability to coordinate with Security/Infrastructure teams and Vendors when security-related issues arise.
Working Skills & Attitude
- Ability to work independently and proactively follow up on issues until they are fully resolved.
- Strong coordination skills when working with Business, Development, Infrastructure, Network, Database teams, and Vendors.
- Ability to work effectively under pressure, particularly during Production Incidents.
- Good communication, reporting, and technical documentation skills.
- Strong willingness to learn and proactively improve processes and automate operational tasks.
- Willingness to participate in on-call/shift support when required.
3. Benefits
- Income VND 20–35 million/month, negotiable based on qualifications and capabilities.
- Annual salary review.
- Seniority benefit: from the 13th month of employment onwards, employees are entitled to an additional 1% seniority bonus for each year of service.
- Excellent career advancement opportunities.
- Accommodation support for employees living far from the workplace.
- Meal allowance/support during work shifts.
- Weekly, monthly, and annual rewards, as well as recognition for valuable contributions to the Company.
- Benefits for special occasions, including visits/support for employees and their families, weddings, funerals, and birthdays.
- Full participation in Social Insurance, Health Insurance, and Unemployment Insurance in accordance with applicable regulations.
Interview & Application Locations
- Hanoi: 98 Son Tay Street, Ba Dinh, Hanoi
- Ho Chi Minh City: 55 Cach Mang Thang 8 Street, Ben Thanh, Ho Chi Minh City
Contact Information
Application Email: tuyendung@a25hotel.com
Recruitment Hotlines:
- Hanoi: 0906 239 925
- Ho Chi Minh City: 0964 468 325