Human Data for RL (Contract)
This job is no longer accepting applications
See open jobs at Bespoke Labs.See open jobs similar to "Human Data for RL (Contract)" 8VC.
Remote
USD 10k-10k / month
Location
Remote
Employment Type
Full time
Location Type
Remote
Department
Human Data
Company Overview
We're an innovative AI startup developing a comprehensive evaluation suite for training and testing autonomous LLM based agents. Our platform pushes the boundaries of AI capability assessment through challenging, real-world problem scenarios.
Position Summary
We're seeking experienced technical professionals to design original, challenging problems that will test the limits of autonomous AI agents across diverse technical domains. This contract role involves creating RL environments that span data science, systems administration, networking, software development, and DevOps scenarios.
Key Responsibilities
Task Design & Development
Design original, hard-to-solve technical problems that challenge autonomous agents across multiple domains.
Create realistic terminal-based scenarios that mirror real-world professional workflows
Develop comprehensive RL environments including problem specifications, execution environments, and grading systems.
Ensure tasks span multiple difficulty levels and technical areas to provide meaningful capability assessment
Domain Coverage
Create tasks across these core technical capabilities:
Data Science and Analysis Workflows: Exploratory Data Analysis, Statistical Hypothesis Testing, A/B Testing, Classification, Regression, Time-Series Forecasting, Clustering & Segmentation, Anomaly Detection
System Administration and Configuration: Server setup, user management, service configuration, log analysis, performance monitoring
Networking and Security Tasks: Network diagnostics, firewall configuration, vulnerability assessment, access control, security auditing
Software Development and Compilation: Code debugging, build systems, dependency management, testing frameworks, deployment pipelines
DevOps and Infrastructure Management: Container orchestration, CI/CD pipelines, infrastructure as code, monitoring and alerting, scaling strategies
Realistic Terminal Scenarios: Command-line problem solving, script automation, file system operations, process management, and other practical terminal-based workflows
Quality Standards
Ensure complete originality - no direct copying from existing LLM benchmarks
Draw inspiration from Kaggle competitions, academic textbooks, MOOCs, and real-world datasets
Pass rigorous plagiarism checks and manual review processes
Create problems that are solvable but genuinely challenging for current AI systems
Required Qualifications
Advanced degree in Data Science, Computer Science, Systems Engineering, or related technical field
3+ years of hands-on experience across multiple technical domains (data science, systems administration, software development, or DevOps)
Strong programming and scripting skills (Python, bash, SQL, and other relevant languages)
Experience with Linux/Unix systems and command-line interfaces
Understanding of containerization, networking, and system configuration
Preferred Qualifications
Background in competitive programming, data science competitions, or capture-the-flag security challenges
Cross-functional expertise spanning data science, DevOps, and systems administration
Experience with cloud platforms, infrastructure as code, and distributed computing
Network security, penetration testing, or cybersecurity experience
Optional: Experience with AI/ML model evaluation and benchmarking
Technical Requirements
Proficiency in creating reproducible, containerized environments
Ability to design fair, robust grading systems that provide meaningful feedback, and checking for reward hacking.
Contract Details
Duration: 2 months with potential for extension
Commitment: Part-time to full-time availability (flexible)
Work Style: Fully remote with weekly office hours and discussions.
Compensation: USD $120/task accepted, with potential to make $10,000+.
Application Process
Please submit:
Resume highlighting relevant data science and evaluation experience
Links to any relevant GitHub repositories, Kaggle profiles, or published work
What We Offer
Opportunity to shape the future of AI evaluation and capability assessment
Flexible remote work environment with autonomous project management
Collaboration with cutting-edge AI research engineers.
Potential for long-term partnership as our platform scales
Next Steps
Screened candidates will be asked to create 1 problem and submit (compensated if they are accepted). If successful, you have the opportunity to submit as many problems as you wish per week. We're looking to fill this position quickly, so early applications are encouraged.
We are an equal opportunity employer committed to diversity and inclusion in our workforce.
This job is no longer accepting applications
See open jobs at Bespoke Labs.See open jobs similar to "Human Data for RL (Contract)" 8VC.