← Tillbaka till jobb

Software Engineer, Machine Learning Platform, New Grad - Quora

  • Distans
  • Sverige
  • Engelska
  • Publicerad 31.08.26 00:04

Tailor my resume to this job

See how your resume can be rewritten and optimized for this specific job. PRO identifies relevant strengths, missing keywords, and opportunities to present your experience more clearly.

Illustrative PRO preview Example output format. Your recommendations will be generated from your resume and this job. Original: Responsible for leading projects and working with different teams.

Suggested rewrite: Led a cross-functional initiative that improved [business outcome] by [measurable result], demonstrating experience relevant to this role...

Am I a good fit for this job?

Compare your resume with this job's requirements. PRO highlights your strongest matches, identifies potential gaps, and recommends practical next steps for your application.

Illustrative PRO preview Example output format. Your actual assessment will be generated from your resume. 1. Overall Match: A personalized match level with an explanation of how your experience aligns with the role.

Key Strengths: Skills, achievements, and experience that support your application. Potential Gaps: Requirements that are missing or not clearly demonstrated in your resume. Recommendation: Specific improvements to strengthen your application...

Cover Letter Assistant

Create a personalized cover letter using this job's requirements and the verified experience in your resume. PRO helps connect your strengths with what the employer is seeking.

Illustrative PRO preview Example output format. Your letter will be personalized using your resume and this job. Dear Hiring Manager, I am excited to apply for this position. My experience in [relevant area] and track record of [measurable achievement] align with the role's key requirements. I would welcome the opportunity to bring these strengths to your team and contribute to [business objective]...

Opportunity details

About This Role.

AI Summary

Quora is hiring a new-graduate Software Engineer to help build and operate its machine-learning platform and ranking infrastructure. The role focuses on distributed systems, reliable model serving, GPU performance optimization, ML developer tooling, and feature-store modernization. Engineers will work with Python, Go, C++, PyTorch, Kubernetes/EKS, NVIDIA Triton, Ray, and AWS while shipping production work early with mentorship from senior engineers. The position includes participation in an on-call rotation and requires overlap with Pacific-time coordination hours. It is well suited to recent or upcoming technical graduates interested in infrastructure and large-scale machine learning systems.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

4/5

EasyHard

Pace & Pressure

4/5

RelaxedFast-paced

Autonomy Level

3/5

GuidedFull ownership

Communication Load

4/5

IndependentCollaborative

AI insightThe role involves production distributed systems, GPU model serving, reliability, and performance work at substantial scale. Although it is designed for new graduates with mentorship, engineers are expected to contribute to production systems quickly and gradually take ownership through on-call participation.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate

$118,300

US market range$95k–$145k

0$160k

AI insightThe disclosed US base-salary range is $97,600 to $139,000 USD annually, with a midpoint of $118,300. This aligns with an estimated US market range of $95,000 to $145,000 annually for a new-graduate software engineer focused on ML infrastructure; equity and benefits are additional and excluded from the base-salary calculation.

Core skills

Skills And Capabilities Most Closely Associated With This Opportunity.

PythonGoC++Machine Learning InfrastructureDistributed SystemsPyTorchKubernetesAWSGPU ServingPerformance Optimization

Sample interview questions

How would you investigate a sudden increase in latency for a production model-serving endpoint?

I would first confirm the scope using latency percentiles, error rates, traffic volume, and resource metrics. I would then isolate whether the bottleneck is in request queuing, model inference, networking, CPU/GPU utilization, memory pressure, or a recent deployment, using traces and profiling data. After identifying the likely cause, I would mitigate impact through rollback, autoscaling, traffic controls, or configuration changes, then validate the fix with monitoring and a documented post-incident review.

What trade-offs would you consider when optimizing GPU model serving?

I would balance latency, throughput, cost, model quality, and reliability. Techniques such as batching, model compilation, precision reduction, concurrency tuning, and instance selection can improve throughput, but may increase tail latency or affect numerical behavior. I would benchmark representative workloads, define service-level objectives, and make changes incrementally with production observability.

Describe how you would design a reliable service for deploying machine-learning models.

I would separate model packaging, validation, deployment, serving, and monitoring into clear stages. The platform should support versioned artifacts, automated compatibility checks, canary releases, rollback, health checks, capacity controls, and metrics for model and infrastructure behavior. I would also prioritize a simple developer workflow so ML engineers can deploy safely without needing to manage all underlying infrastructure details.

How would you approach learning an unfamiliar infrastructure technology such as Kubernetes or NVIDIA Triton?

I would begin with the core concepts and build a small hands-on project that exercises the most relevant workflow. Next, I would read existing internal configurations and documentation, ask targeted questions of experienced teammates, and use tests or a sandbox environment to validate my understanding. I would document what I learn and gradually take on well-scoped production tasks with appropriate review.

What makes an effective on-call engineer, especially for someone early in their career?

An effective on-call engineer stays calm, communicates status clearly, follows established runbooks, and focuses first on reducing user impact. For a new engineer, I would prepare by understanding alerts, dashboards, escalation paths, and common failure modes, while pairing with experienced teammates when needed. After an incident, I would help improve monitoring, documentation, or automation so the same issue is easier to handle in the future.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

[Quora is a privately held, “remote-first” company. This position can be performed remotely from anywhere in Canada or the United States. Please visit careers.quora.com/eligible-countries for details regarding employment eligibility by country.]

About Quora

Quora’s mission is to grow the world’s collective intelligence. To do so, we have two platforms:

Quora: a global knowledge sharing platform with over 300M monthly unique visitors, bringing people together to share insights on various topics and providing a unique platform to learn and connect with others.Poe: a platform providing millions of global users with one place to chat, explore and build with a wide variety of AI language models (bots), including GPT-5.6-Sol, Claude-Opus-5, Claude-Fable-5, Claude-Sonnet-5, Kimi-K3, and thousands of others. As AI capabilities rapidly advance, Poe provides a single platform to instantly integrate and utilize these new models.

Behind these products are passionate, collaborative, and high-performing global teams. We have a culture rooted in transparency, idea-sharing, and experimentation that allows us to celebrate success and grow together through meaningful work. Join us on this journey to create a positive impact and make a significant change in the world.

This role will be working on our Quora product.

About The Team And Role

Machine Learning is central to Quora’s mission of growing the world’s collective intelligence. We have 100+ Machine Learning models in production powering various product features. We use a variety of algorithms — everything from linear models to decision trees and deep neural networks. Our production models operate at a huge scale, serving hundreds of millions of people using Quora every month.

Our team owns Quora’s ML platform and ranking infrastructure across four areas: serving reliability, ML engineer enablement and developer velocity, business impact, and cost efficiency. We want to empower all ML engineers at Quora to be as impactful as they can be in solving different ML problems at scale.

As a Software Engineer (New Grad) on this team, you’ll work at the intersection of Machine Learning, Distributed Systems, and GPU Serving performance — and your work will have an enormous impact on Quora’s long-term success.

No previous ML infrastructure experience is required for this role. You’ll be joining a team of senior and staff engineers, learning this stack from the people who built it, with a dedicated mentor and strong technical guidance — and you’ll be shipping to production in your first few weeks.

Stack: Python, Go, C++, PyTorch, Kubernetes/EKS, NVIDIA Triton, Ray, AWS

🚀 Excited to see our MLP team’s amazing work in action? Check out some of the incredible projects they’ve completed below! 👇✨

– https://quoraengineering.quora.com/Migrating-from-x86-to-AWS-Graviton-A-Journey-in-Cost-Optimization-and-Performance

– https://aws.amazon.com/blogs/containers/quora-3x-faster-machine-learning-25-lower-costs-with-nvidia-triton-on-amazon-eks/

– https://quoraengineering.quora.com/Building-a-Service-Mesh-in-a-Hybrid-Environment

– https://quoraengineering.quora.com/Building-Embedding-Search-at-Quora

– https://quoraengineering.quora.com/Feature-Engineering-at-Quora-with-Alchemy

Responsibilities

Help build and maintain the core infrastructure that powers Quora’s ML platform, ensuring high availability, scalability, and performanceBuild and improve the distributed systems that serve our ML models in production, from Large Recommendation Models (LRM) to Large Language Models (LLM)Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable modelsContribute to platform initiatives such as PyTorch-first standardization and ML ecosystem modernizationImprove ML developer velocity by building tooling that helps ML engineers develop, test, and deploy models more efficientlyModernize our feature store so ML engineers can get new features into production fasterParticipate in the team’s on-call rotation, helping resolve production issues as you grow your knowledge and ownership of the platform

Minimum Requirements

Availability for meetings and impromptu communication during Quora’s “coordination hours” (Mon-Fri: 9am-3pm Pacific Time)A 2025 or 2026 graduate with or pursuing a B.S., M.S., or Ph.D. in Computer Science, Engineering or a related technical fieldGenuine interest in large-scale distributed systems, infrastructure, and machine learningKnowledge of Python, Go or C++, or the ability to learn them quicklyA passion for learning and always improving yourself and the team around you

Preferred Requirements

Previous software engineering experience via an internship, work experience, open-source contribution or coding competitionCoursework or hands-on experience with ML frameworks such as PyTorch or TensorFlowExposure to Kubernetes, Docker, or cloud technologies like AWSExperience with low-level performance work of any kind: profiling, benchmarking, optimizationPassion for Quora’s mission and goals

At Quora, we value diversity and inclusivity and welcome individuals from all backgrounds, including marginalized or underrepresented groups in tech, to apply for our job openings. We encourage all candidates who share a passion for growing the world’s knowledge, even those who may not strictly meet all the preferred requirements, to apply, as we know that a diverse range of perspectives can have a significant impact on our products and our culture.

Additional Information

We are accepting applications on an ongoing basis. This role is a backfill for an existing vacancy.

Quora offers a wide range of benefits including medical/dental/vision coverage, equity refreshers, remote work reimbursement, paid time off, employee assistance programs, and more. Benefits are country-specific and may vary.

There are many factors that will determine the starting pay, including but not limited to experience, location, education, and business needs.

US candidates only: For US based applicants, the salary range is $97,600 – $139,000 USD + equity + benefits.Canada candidates only: For Toronto and Vancouver based applicants, the salary range is $125,320 – $142,783 CAD + equity + benefits. For all other locations in Canada, the salary range is $116,965 – $133,264 CAD + equity + benefits.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

AI technology may assist in sorting applications and recording interview notes, but all decisions are made by a member of our team.

To ensure a secure hiring process, all final candidates will undergo identity verification and a comprehensive background check prior to onboarding.

Job Applicant Privacy Notice: https://www.careers.quora.com/pages/quora-global-job-applicant-privacy-notice

Apply now >

Upload your resume nowTo unlock remote work opportunities and be discovered by global employers. This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Job Complexity

4/5

EasyHard

Pace & Pressure

4/5

RelaxedFast-paced

Autonomy Level

3/5

GuidedFull ownership

Communication Load

4/5

IndependentCollaborative

Estimated job medianMarket rate

$118,300US market range$95k–$145kHow would you investigate a sudden increase in latency for a production model-serving endpoint?

I would first confirm the scope using latency percentiles, error rates, traffic volume, and resource metrics. I would then isolate whether the bottleneck is in request queuing, model inference, networking, CPU/GPU utilization, memory pressure, or a recent deployment, using traces and profiling data. After identifying the likely cause, I would mitigate impact through rollback, autoscaling, traffic controls, or configuration changes, then validate the fix with monitoring and a documented post-incident review.

What trade-offs would you consider when optimizing GPU model serving?

I would balance latency, throughput, cost, model quality, and reliability. Techniques such as batching, model compilation, precision reduction, concurrency tuning, and instance selection can improve throughput, but may increase tail latency or affect numerical behavior. I would benchmark representative workloads, define service-level objectives, and make changes incrementally with production observability.

Describe how you would design a reliable service for deploying machine-learning models.

I would separate model packaging, validation, deployment, serving, and monitoring into clear stages. The platform should support versioned artifacts, automated compatibility checks, canary releases, rollback, health checks, capacity controls, and metrics for model and infrastructure behavior. I would also prioritize a simple developer workflow so ML engineers can deploy safely without needing to manage all underlying infrastructure details.

How would you approach learning an unfamiliar infrastructure technology such as Kubernetes or NVIDIA Triton?

I would begin with the core concepts and build a small hands-on project that exercises the most relevant workflow. Next, I would read existing internal configurations and documentation, ask targeted questions of experienced teammates, and use tests or a sandbox environment to validate my understanding. I would document what I learn and gradually take on well-scoped production tasks with appropriate review.

What makes an effective on-call engineer, especially for someone early in their career?

An effective on-call engineer stays calm, communicates status clearly, follows established runbooks, and focuses first on reducing user impact. For a new engineer, I would prepare by understanding alerts, dashboards, escalation paths, and common failure modes, while pairing with experienced teammates when needed. After an incident, I would help improve monitoring, documentation, or automation so the same issue is easier to handle in the future.

Axon Aug 30

100524 – Senior Software Engineer II

Join Axon and be a Force for Good. At Axon, we’re on a mission to Protect Life. We’re explorers, pursuing society’s most critical safety and justice issues with our ecosystem…

New

Remote from

US

Work type

Full Time

Compensation

USD 148,500–237,600/year

Join Axon and be a Force for Good. At Axon, we’re on a mission to Protect Life. We’re explorers, pursuing society’s most critical safety and justice issues with our ecosystem…

Testlio Aug 30

Freelance Software Tester with DirecTV Satellite Dish

Hi there! We are Testlio, a global software testing company with its own freelance network. Our freelancers test apps from companies across the globe. Testlio is a great place for…

New

Remote from

US

Work type

Contract

Compensation

USD 22–22/hour

Hi there! We are Testlio, a global software testing company with its own freelance network. Our freelancers test apps from companies across the globe. Testlio is a great place for…

Teachable Aug 30

Senior Full Stack Engineer II

What is Teachable? Teachable is the platform for experts and businesses who take education seriously. In a world where anyone can ask AI for information, we’re the home for those…

New

Remote from

BR

Work type

Full Time

What is Teachable? Teachable is the platform for experts and businesses who take education seriously. In a world where anyone can ask AI for information, we’re the home for those…

Testlio Aug 30

AI Solutions Architect (United States)

Location: This is a 100% remote role open to those who reside in the United States. About the job Testlio’s fully managed crowdsourced testing platform, powered by our proprietary intelligence…

New

Remote from

US

Work type

Full Time

Location: This is a 100% remote role open to those who reside in the United States. About the job Testlio’s fully managed crowdsourced testing platform, powered by our proprietary intelligence…

Teachable Aug 30

Engineering Manager

What is Teachable? Teachable is the platform for experts and businesses who take education seriously. In a world where anyone can ask AI for information, we’re the home for those…

New

Remote from

BR

Work type

Full Time

What is Teachable? Teachable is the platform for experts and businesses who take education seriously. In a world where anyone can ask AI for information, we’re the home for those…

Windranger Labs Aug 30

QA Engineer 区块链测试工程师

Who we are Mantle is the entry point for institutions and traditional finance to access real-world assets on-chain. Backed by the world’s largest community-owned treasury of over $4B, Mantle combines…

New

Remote from

APAC

Work type

Full Time

Who we are Mantle is the entry point for institutions and traditional finance to access real-world assets on-chain. Backed by the world’s largest community-owned treasury of over $4B, Mantle combines…

Windranger Labs Aug 30

Blockchain Engineer 区块链开发工程师

Who we are Mantle is the entry point for institutions and traditional finance to access real-world assets on-chain. Backed by the world’s largest community-owned treasury of over $4B, Mantle combines…

New

Remote from

APAC

Work type

Full Time

Who we are Mantle is the entry point for institutions and traditional finance to access real-world assets on-chain. Backed by the world’s largest community-owned treasury of over $4B, Mantle combines…

UpGuard Aug 30

Software Engineer (Multiple Levels)

Who are we?At UpGuard, we are replacing manual security bottlenecks with AI-driven precision. Fresh off a US$75M Series C, we are scaling our infrastructure to process 100 billion risk signals…

New

Remote from

AU

Work type

Full Time

Who are we?At UpGuard, we are replacing manual security bottlenecks with AI-driven precision. Fresh off a US$75M Series C, we are scaling our infrastructure to process 100 billion risk signals…

Reddit Aug 30

Staff Software Engineer, Media Experiences

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit…

New

Remote from

US

Work type

Full Time

Compensation

USD 217k–303k/year

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit…

Supabase Aug 29

Developer Relations Engineer

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All…

New

Remote from

US

Work type

Full Time

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All…