Salesforce · Behavioral
AI / ML Fundamentals Oral Round (AI Engineer)
TrueInterview
September 26, 2026 · 3 min read
Requirements
- This is a spoken-only interview; you won't be asked to code on a whiteboard (though you should be prepared to sketch on a shared board if the interviewer requests it).
- Questions fall into three broad categories:
- ML fundamentals: the difference between classification and regression, F1 versus precision versus recall, the meaning of F1 and when to favor it, benchmarking practices, proper train/validation/test splitting, and typical loss functions.
- Model serving and pipelines: inference optimization techniques (quantization, batching, KV-cache, latency-throughput trade-offs), and how to deploy an AI agent as a service (state handling, idempotency, observability, rollback).
- LLM-focused topics: the Transformer architecture, intuition behind the attention mechanism, context engineering versus prompting, retrieval-augmented generation (RAG) architecture, grounding (factuality and source attribution), and guardrails (content filters, jailbreak resistance, output validation).
Notes
- This round values fluency with concepts more than deep expertise. Interviewers would rather hear a candidate connect three related ideas in a coherent chain (for instance, "context length scales quadratically with attention computation; we cache the KV state across tokens to amortize the cost; sliding-window attention limits the active context") than give single-word responses.
- F1-score: the harmonic mean of precision and recall, given by . It is preferred over accuracy when classes are imbalanced. Be familiar with macro, micro, and weighted variants.
- Classification vs. regression: discrete versus continuous outputs; different loss functions (cross-entropy vs. mean squared error); different evaluation metrics. Be prepared to provide one example of each from a domain relevant to Salesforce (lead scoring, churn prediction, sentiment analysis).
- Benchmarking: use a held-out test set that matches the production distribution; monitor multiple metrics, not just one; verify there is no data leakage between training and test sets; report confidence intervals when the sample size is small.
- Inference optimization: batching (balancing latency and throughput), KV-cache (to avoid recomputing previous attention), quantization (FP16, INT8, INT4 — trading off accuracy for speed), speculative decoding, and prompt caching.
- Deploying an AI agent as a service: stateful (conversation history) versus stateless (per-turn) request models; idempotency for retried tool calls; observability (per-turn tracing, tool-call auditing); rollback strategies (versioned prompts, A/B routing).
- Transformer: attention(Q, K, V) = softmax(QK^T / √d_k) V; multi-head attention splits the channel dimension; positional encoding (sinusoidal, learned, or RoPE); the decoder applies a causal mask. Understand the relationship among d_model, d_k, and h.
- RAG: a retriever (sparse BM25, dense embeddings, or hybrid) → top-k chunks → inserted into the prompt → the LLM generates with citations. Know the trade-offs of chunk size, the role of reranking, and evaluation methods (faithfulness, relevance, answer correctness).
- Grounding: linking outputs to the retrieved evidence; reduces hallucination; typically implemented via citations and source attribution in the response.
- Guardrails: input filters (prompt injection detection, PII scrubbing), output filters (toxicity, jailbreak detection, format validation), and retrieval gating. Be able to name one open-source library (such as NeMo Guardrails, Guardrails AI, or Llama Guard).
Preparation
- Create a one-page cheat sheet covering all three categories and practice delivering a 60–90 second explanation for each key concept.
- For LLMs in particular, review: the attention formula, KV-cache mechanics, end-to-end RAG, the distinction between grounding and guardrails, and evaluation metrics (faithfulness, helpfulness, harmlessness).
- Practice walking through the inference optimization stack from scratch ("raw model → quantized → batched → KV-cached → speculative-decoded") — this question often separates prepared candidates from unprepared ones.
- Read a recent production LLM serving guide and extract a concrete checklist covering latency, throughput, batching, and rollback.
Loading comments…