SDET + Python + Playwright + Document Processing (Xbrl)

Worldwide | Sept. 19, 2026

Report as Closed

Company: Improving

Country: Worldwide

Salary: $42,000 - $54,000

Type: Remote

Employment: Full-time

Description: Improving is an IT services firm focused on AI, Data, and Applications. We modernize legacy systems, build cloud-native platforms, and deliver future-ready solutions through collaborative, long-term partnerships that drive measurable outcomes.

At Improving South America, we provide IT services to transform the perception of the IT professional. We focus on IT consulting, software development and agile training. 
The company promotes an exceptional work culture based on teamwork, excellence and fun, with a focus on personal growth and shared rewards. By joining, the candidate will become part of a community that prioritizes open communication and strong long-term working relationships, supported by a structure of professional development and continuous learning.
We are looking for a strong technical engineer who can design evaluation strategies, implement validation frameworks, and turn regulatory requirements into automated, measurable quality gates.
  • Strong technical background validating and testing complex systems, with hands-on experience building test harnesses.
  • Experience building automated test frameworks, including creation of golden datasets, custom validation frameworks, and automated correctness/validation logic.
  • Proven experience designing and implementing benchmarking strategies and quality metrics for complex outputs.
  • Familiarity with LLM evaluation, validation techniques, and quality metrics (including how to interpret results and drive iteration).
  • Proficiency with test automation tools, with Playwright preferred for web/HTML-level validation.
  • Solid understanding of XBRL, regulations.
  • Background in data quality validation and metrics design.
  • Experience with CI/CD pipelines and DevOps practices to run benchmarks and validations consistently.
Soft skills we value: we expect clear written communication, structured thinking, ownership of quality outcomes, and an iterative mindset that balances compliance precision with practical engineering constraints.
In this role, we are focused on production readiness for AI-assisted financial/regulatory tagging. You will design and implement a comprehensive benchmarking framework that measures Large Language Model (LLM) put quality across multiple dimensions, then establish the test automation and scoring mechanisms needed to track improvements over time. You will work closely with our R&D team on the AI tagging core and with production engineering to identify gaps, iterate on quality gates, and ensure the system can reliably meet regulatory and business requirements.
Responsibilities
We will have you lead the design and implementation of a benchmarking and validation framework for LLM outputs, with clear, measurable quality dimensions and automated checks.
  • Design and implement a comprehensive benchmarking framework that measures LLM output quality across multiple dimensions, including:
  • Technical correctness, such as valid XHTML and IXBRL schema compliance.
  • SEC validation compliance, ensuring generated XBRL instance documents pass SEC validation without errors.
  • Completeness, verifying all required facts are tagged (including hidden or non-obvious data points).
  • Create complexity “tranches” for benchmarking, ranging from single fund / single share class through multi-fund scenarios and 200+ share class cases, and define success metrics for each tier.
  • Build test automation for web and HTML-level testing using Playwright (with no existing automation today), including the first stable automated test harness.
  • Create measurable progress indicators that track improvements as the AI tagging core is enhanced and refined.
  • Validate and score LLM outputs against both regulatory and business requirements, producing repeatable results suitable for ongoing quality evaluation.
  • Collaborate with the R&D team (focused on AI tagging) and production engineering to identify gaps, prioritize iteration steps, and support production readiness.
Benefits
  • Long-term contract.
  • 100% Remote.
  • Vacation and PTOs
  • Apply here:

    Web: Apply here

    Emails:



Similar Python Jobs

Found 6 similar Remote jobs

Remote Full-time

Worldwide

$18,000 - $30,000

View Job →
Onboarding Operations Tech Lead

Twilio

Remote Full-time

Worldwide

View Job →
Software Engineer, Ingestion Platform

Reddit

Remote Full-time

Worldwide

View Job →
Staff Product Operations Lead, Marketplace Operations

Databricks

Remote Full-time

Worldwide

View Job →
Sr. Solutions Engineer - Digital Native Business (Strategic …

Databricks

Remote Full-time

Worldwide

View Job →
Bluelight Consulting
Senior Python/Flask+React Fullstack Engineer - Remote, Latin America

Bluelight Consulting

Remote Full Time

Worldwide

View Job →

Find High-Paying Python Developer Jobs ($80K-$200K+)

Django · Flask · FastAPI · Remote & Onsite · Updated daily · No recruiter spam
💼 Get the best Python jobs weekly. Salary-transparent roles, no recruiter spam — unsubscribe anytime.