ML Engineer / Data Scientist
Worldwide | Sept. 1, 2026
Report as Closed
Company: Lisit
Country: Worldwide
Type: Remote
Employment: Full-time
Description: We create, develop, and implement software services that provide the necessary automation and optimization tools, while maintaining a constant focus on innovation and a passion for challenges.
At Lisit, we design, develop, and implement software services focused on process automation and optimization. Our goal is to offer tools that enable clients to transform their operations with a constant focus on innovation and a passion for challenges. The team promotes operational efficiency through consultative support that integrates diverse tools and practices to achieve their objectives, using a comprehensive support and implementation strategy.
Required
- Strong Python skills (pandas, scikit-learn) and real-world experience taking an NLP model in Spanish to production.
- Experience with tokenization, text normalization, and handling spelling errors.
- Experience with Transformers / Hugging Face for BERT fine-tuning (ideally BETO or another Spanish-language BERT) focused on NER.
- Programmatic labeling / weak supervision: building training datasets without massive manual labeling.
- Rigorous evaluation: precision/recall/F1, held-out sets, and data leakage control.
How we evaluate the approach: We need you to clearly explain how you would build the training dataset when there is no labeled corpus and there are 1.8 million historical pairs; and what it means for a model to be calibrated and how you would verify it.
We also value an organized and communicative work style: a focus on reproducibility, critical thinking about metrics (held-out sets and data leakage), and a customer-centric mindset.
At Lisit, we create, develop, and implement software services with a constant focus on innovation and a passion for challenges. In this role, we work on automating and optimizing processes through Spanish NLP solutions that have a direct impact on our clients’ operational efficiency. The project’s mission is to build and evolve a reproducible pipeline based on
1.8 million historical pairs, transforming it into a training dataset, train and evaluate a Spanish NER model, calibrate the confidence score, and facilitate the transfer of knowledge and the solution to the client for sustained adoption.
Mission
Mission: build the training corpus from 1.8 million historical pairs; fine-tune a Spanish NER model; calibrate the confidence score; and ensure the pipeline is reproducible and transferable to the client.
- Build the training dataset without extensive manual labeling (programmatic labeling / weak supervision).
- Perform text tokenization and normalization for NLP in Spanish.
- Train and fine-tune a Spanish NER model (ideally Spanish BERT, such as BETO or another Spanish BERT) using Transformers / Hugging Face.
- Handle spelling errors and ensure robust preprocessing.
- Rigorously evaluate performance using precision, recall, and F1 score with held-out sets and data leakage control.
- Calibrate probabilities so that the confidence threshold is interpretable (score calibrated using methods such as Platt or isotonic calibration).
- Ensure reproducibility: document the pipeline, ensure experiment traceability, and package it so the client can replicate and operate it.
Benefits
100% remote
If you’re interested in building a reproducible pipeline, training NER in Spanish with a focus on weak supervision, and rigorously evaluating and calibrating the model, we invite you to apply.
Desirable (bonus points)
- XGBoost and probability calibration (Platt / isotonic): the threshold of 0.82 must correspond to the true probability.
- Embeddings and vector search (pgvector, HNSW, sentence-transformers).
- Fuzzy matching: edit distances, phonetic algorithms, pg_trgm.
- MLflow for experiment management and deployment.
- ONNX and INT8 quantization for efficient inference on the CPU (without a GPU).
- libpostal.
- Experience with mailing addresses, geocoding, or geographic data.
Apply here:
Web:
Apply here
Emails: