Human Data for RL (Contract)

Bespoke Labs
Bespoke Labs

Remote

USD 10k-10k / month

Posted 6+ months ago

Location

Remote

Employment Type

Full time

Location Type

Remote

Department

Human Data

Company Overview

We're an innovative AI startup developing a comprehensive evaluation suite for training and testing autonomous LLM based agents. Our platform pushes the boundaries of AI capability assessment through challenging, real-world problem scenarios.

Position Summary

We're seeking experienced technical professionals to design original, challenging problems that will test the limits of autonomous AI agents across diverse technical domains. This contract role involves creating RL environments that span data science, systems administration, networking, software development, and DevOps scenarios.

Key Responsibilities

Task Design & Development

  • Design original, hard-to-solve technical problems that challenge autonomous agents across multiple domains.

  • Create realistic terminal-based scenarios that mirror real-world professional workflows

  • Develop comprehensive RL environments including problem specifications, execution environments, and grading systems.

  • Ensure tasks span multiple difficulty levels and technical areas to provide meaningful capability assessment

Domain Coverage

Create tasks across these core technical capabilities:

  • Data Science and Analysis Workflows: Exploratory Data Analysis, Statistical Hypothesis Testing, A/B Testing, Classification, Regression, Time-Series Forecasting, Clustering & Segmentation, Anomaly Detection

  • System Administration and Configuration: Server setup, user management, service configuration, log analysis, performance monitoring

  • Networking and Security Tasks: Network diagnostics, firewall configuration, vulnerability assessment, access control, security auditing

  • Software Development and Compilation: Code debugging, build systems, dependency management, testing frameworks, deployment pipelines

  • DevOps and Infrastructure Management: Container orchestration, CI/CD pipelines, infrastructure as code, monitoring and alerting, scaling strategies

  • Realistic Terminal Scenarios: Command-line problem solving, script automation, file system operations, process management, and other practical terminal-based workflows

Quality Standards

  • Ensure complete originality - no direct copying from existing LLM benchmarks

  • Draw inspiration from Kaggle competitions, academic textbooks, MOOCs, and real-world datasets

  • Pass rigorous plagiarism checks and manual review processes

  • Create problems that are solvable but genuinely challenging for current AI systems

Required Qualifications

  • Advanced degree in Data Science, Computer Science, Systems Engineering, or related technical field

  • 3+ years of hands-on experience across multiple technical domains (data science, systems administration, software development, or DevOps)

  • Strong programming and scripting skills (Python, bash, SQL, and other relevant languages)

  • Experience with Linux/Unix systems and command-line interfaces

  • Understanding of containerization, networking, and system configuration

Preferred Qualifications

  • Background in competitive programming, data science competitions, or capture-the-flag security challenges

  • Cross-functional expertise spanning data science, DevOps, and systems administration

  • Experience with cloud platforms, infrastructure as code, and distributed computing

  • Network security, penetration testing, or cybersecurity experience

  • Optional: Experience with AI/ML model evaluation and benchmarking

Technical Requirements

  • Proficiency in creating reproducible, containerized environments

  • Ability to design fair, robust grading systems that provide meaningful feedback, and checking for reward hacking.

Contract Details

  • Duration: 2 months with potential for extension

  • Commitment: Part-time to full-time availability (flexible)

  • Work Style: Fully remote with weekly office hours and discussions.

  • Compensation: USD $120/task accepted, with potential to make $10,000+.

Application Process

Please submit:

  1. Resume highlighting relevant data science and evaluation experience

  2. Links to any relevant GitHub repositories, Kaggle profiles, or published work

What We Offer

  • Opportunity to shape the future of AI evaluation and capability assessment

  • Flexible remote work environment with autonomous project management

  • Collaboration with cutting-edge AI research engineers.

  • Potential for long-term partnership as our platform scales

Next Steps

Screened candidates will be asked to create 1 problem and submit (compensated if they are accepted). If successful, you have the opportunity to submit as many problems as you wish per week. We're looking to fill this position quickly, so early applications are encouraged.

We are an equal opportunity employer committed to diversity and inclusion in our workforce.