xelys jobs xelys jobs

Mobile Software Engineer – Swift & AI Evaluation | Remote

Crossing Hurdles

full-remoteseniorcontractbackend United States 180 days ago via LinkedIn
72,000 - 120,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

SwiftLLMsAI EvaluationOpen SourceCode ReviewTechnical WritingFact-CheckingSoftware EngineeringBenchmarkingAlgorithmic Soundness

About the role

Role Overview

Mobile Software Engineer – Swift & AI Evaluation (Hourly contract, remote) You will evaluate AI-generated responses related to Swift and software engineering, validating correctness and providing structured technical feedback.

Responsibilities

  • Evaluate AI-generated responses for coding/software engineering queries (accuracy, reasoning, clarity, completeness)
  • Fact-check using reliable public sources and references
  • Execute code to validate correctness and test outputs
  • Annotate responses by identifying strengths, issues, and inaccuracies
  • Assess code quality, including readability and algorithmic soundness
  • Ensure responses align with best practices and expected system behavior
  • Follow structured evaluation guidelines, benchmarks, and taxonomies

Requirements

  • Strong software engineering background in related technical roles
  • Strong expertise in Swift
  • Ability to solve medium to hard coding problems independently
  • Experience contributing to open-source projects
  • Experience using LLMs for coding and understanding their limitations
  • High attention to detail when evaluating technical reasoning
  • Clear written communication for technical feedback
  • Bachelor’s degree in Computer Science or a related field

Application Process

  • Upload resume
  • Complete a brief interview
  • Submit a short form

About Crossing Hurdles

Crossing Hurdles is an organization focused on technology evaluation and applied data science work. The role involves assessing AI-generated coding responses and ensuring they meet accuracy and best-practice standards.

Scraped 4/1/2026