System Design · Jane Street · Hard
You are given DNA strings over the alphabet {A, C, G, T}. The strings are not all the same length. You have access to a small annotated dataset where each sequence is labeled as genuine or fake, and a much larger unannotated dataset containing only sequences known to be genuine. Your task is to design a machine learning system that predicts whether a previously unseen DNA sequence is genuine or fake. Your response should describe: Input representation: How would you encode…
Checking your access…