Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
NLPEasy
Q1. What does NLP stand for?
- A.Numerical Linear Processing
- B.Network Logic Protocol
- C.Natural Language Processing✓ Correct
- D.Neural Language Programming
Explanation
NLP stands for Natural Language Processing, the field of AI dealing with interaction between computers and human language.
Report an error in this question
NLPEasy
Q2. Named Entity Recognition (NER) identifies:
- A.Grammar errors found throughout the given text✓ Correct
- B.Spelling mistakes identified within the text body
- C.The overall length of each sentence in the document
- D.Named entities like people, organizations, and locations in text
Explanation
NER identifies and classifies named entities in text into predefined categories such as person, organization, location, and date.
Report an error in this question
NLPEasy
Q3. What is stemming?
- A.Reducing words to their root form by removing suffixes✓ Correct
- B.Adding prefixes to words for morphological expansion
- C.Counting the total frequency of words in text
- D.Translating words between two different languages
Explanation
Stemming removes suffixes to reduce words to their root form (e.g., 'running' -> 'run'), simplifying text for processing.
Report an error in this question
NLPEasy
Q4. What is a corpus in NLP?
- A.A large collection of text documents✓ Correct
- B.A specific grammar rule definition
- C.A single word in a sentence
- D.A type of neural network model
Explanation
A corpus is a large, structured collection of text documents used for training and evaluating NLP models.
Report an error in this question
NLPEasy
Q5. What is lemmatization?
- A.Counting the total characters in a word
- B.Removing all vowels from the input text
- C.Adding suffixes for morphological inflection
- D.Reducing words to their base dictionary form✓ Correct
Explanation
Lemmatization reduces words to their base dictionary form (lemma) using vocabulary and morphological analysis (e.g., 'better' -> 'good').
Report an error in this question
NLPEasy
Q6. Tokenization in NLP means:
- A.Encrypting text for secure data transmission
- B.Splitting text into individual words or subwords✓ Correct
- C.Deleting text from the data pipeline
- D.Compressing text to reduce storage size
Explanation
Tokenization breaks text into smaller units (tokens) such as words, subwords, or characters for processing by NLP models.
Report an error in this question
NLPEasy
Q7. Stop words are:
- A.Important domain-specific keywords carrying high information
- B.Common words like 'the' and 'is' often removed in processing✓ Correct
- C.Rare words that appear only once in the entire corpus
- D.Technical terms from a specialized subject vocabulary
Explanation
Stop words are frequently occurring words that carry little meaning and are often removed during text preprocessing to focus on content words.
Report an error in this question
NLPEasy
Q8. What is sentiment analysis?
- A.Summarizing long documents into shorter text passages
- B.Translating text from one natural language to another language
- C.Generating entirely new text from a given prompt input
- D.Determining the emotional tone of text (positive, negative, neutral)✓ Correct
Explanation
Sentiment analysis identifies the emotional tone or opinion expressed in text, classifying it as positive, negative, or neutral.
Report an error in this question
NLPEasy
Q9. What is text classification?
- A.Generating entirely new text from a prompt
- B.Translating text between different languages
- C.Editing and correcting text for grammar
- D.Assigning predefined categories to text documents✓ Correct
Explanation
Text classification assigns predefined labels or categories to text documents, such as spam detection or topic classification.
Report an error in this question
NLPEasy
Q10. What is machine translation?
- A.Performing manual text translation
- B.Generating new text from a prompt
- C.Summarizing text into shorter form✓ Correct
- D.Automatically translating text from one language to another
Explanation
Machine translation uses AI to automatically translate text between languages, as in Google Translate.
Report an error in this question
NLPMedium
Q11. What is the bag-of-words model?
- A.Representing text with full word order preserved
- B.A syntactic parsing technique for sentences
- C.Representing text as the frequency of words, ignoring order✓ Correct
- D.A neural language model using word embeddings
Explanation
Bag-of-words represents text as a vector of word frequencies, disregarding grammar and word order while capturing word presence.
Report an error in this question
NLPMedium
Q12. What is Word2Vec?
- A.A simple word frequency counter for bag-of-words features
- B.A technique learning dense vector representations of words from context✓ Correct
- C.A language dictionary for looking up word definitions
- D.A rule-based grammar checker for text correction
Explanation
Word2Vec learns dense vector representations (embeddings) of words where semantically similar words are mapped to nearby points in vector space.
Report an error in this question
NLPMedium
Q13. What is attention in sequence-to-sequence models?
- A.A data augmentation method for creating synthetic text training data
- B.A mechanism allowing the decoder to focus on relevant parts of the input sequence✓ Correct
- C.A cross-entropy loss function for measuring sequence prediction accuracy
- D.A regularization technique for preventing overfitting in sequence models
Explanation
Attention allows the decoder to attend to different parts of the input sequence when generating each output token, improving translation and generation quality.
Report an error in this question
NLPMedium
Q14. What is TF-IDF?
- A.A recurrent neural network variant designed for sequential text modeling
- B.A structured query language for relational database data retrieval
- C.A multi-layer deep feedforward neural network architecture for classification
- D.A weighting scheme reflecting word importance in a document relative to a corpus✓ Correct
Explanation
TF-IDF (Term Frequency-Inverse Document Frequency) weighs words by how often they appear in a document relative to how common they are across all documents.
Report an error in this question
NLPMedium
Q15. What is text embedding?
- A.Converting text into dense numerical vectors✓ Correct
- B.Removing text from the raw dataset
- C.Encrypting text for secure storage
- D.Converting numbers into readable text
Explanation
Text embedding converts text into dense numerical vector representations that capture semantic meaning, enabling mathematical operations on text.
Report an error in this question
NLPMedium
Q16. What is sequence-to-sequence (Seq2Seq) modeling?
- A.Mapping a single input word to a single corresponding output word
- B.Sorting input sequences into a predefined correct order
- C.Clustering multiple sequences into meaningful related groups✓ Correct
- D.Mapping an input sequence to an output sequence of potentially different length
Explanation
Seq2Seq models transform an input sequence into an output sequence, used in translation, summarization, and dialogue generation.
Report an error in this question
NLPMedium
Q17. What is BERT?
- A.A rule-based expert system for language processing
- B.A generative adversarial network designed for text data
- C.A pre-trained transformer for bidirectional language understanding✓ Correct
- D.A type of recurrent neural network for sequential processing
Explanation
BERT (Bidirectional Encoder Representations from Transformers) is pre-trained on masked language modeling and next sentence prediction for bidirectional context understanding.
Report an error in this question
NLPMedium
Q18. What is part-of-speech (POS) tagging?
- A.Removing individual words from the input text sequence
- B.Translating text between two different natural languages
- C.Assigning grammatical categories (noun, verb, etc.) to each word✓ Correct
- D.Counting the number of syllables in each word
Explanation
POS tagging assigns grammatical labels like noun, verb, adjective to each word in a sentence, aiding in syntactic analysis.
Report an error in this question
NLPHard
Q19. What is the difference between GPT and BERT architectures?
- A.Both use exactly the same attention mechanism and training objective
- B.GPT uses bidirectional attention; BERT uses causal left-to-right attention
- C.They are completely identical transformer architectures with no differences
- D.GPT uses causal attention for generation; BERT uses bidirectional attention for understanding✓ Correct
Explanation
GPT uses masked (causal) self-attention for autoregressive text generation. BERT uses bidirectional self-attention for understanding tasks, seeing context from both directions.
Report an error in this question
NLPHard
Q20. What is the perplexity metric in language models?
- A.The total elapsed time required for training the model
- B.The number of parameters or total size of the model
- C.A measure of how well a model predicts a sample; lower is better✓ Correct
- D.The total number of words in the vocabulary of the model
Explanation
Perplexity measures how well a probability distribution predicts a sample. Lower perplexity indicates the model is better at predicting the text.
Report an error in this question
NLPMedium
Q21. What is cosine similarity used for in NLP?
- A.Counting the total frequency of words in text
- B.Parsing sentences into syntax tree structures
- C.Generating new text from a trained model
- D.Measuring the similarity between two text vectors✓ Correct
Explanation
Cosine similarity measures the cosine of the angle between two vectors, indicating how similar two texts are regardless of their magnitude.
Report an error in this question
NLPHard
Q22. What is byte-pair encoding (BPE) in tokenization?
- A.A tokenization approach that only works at the word level
- B.A tokenization strategy that works at sentence level✓ Correct
- C.A tokenization method operating at character level only
- D.A subword tokenization algorithm that iteratively merges the most frequent character pairs
Explanation
BPE starts with characters and iteratively merges the most frequent adjacent pairs to create subword units, balancing vocabulary size and out-of-vocabulary handling.
Report an error in this question
NLPHard
Q23. What is cross-lingual transfer learning?
- A.Translating all text first then training on the translated version
- B.Ignoring language differences and treating all text the same
- C.Training on one language and applying knowledge to another language✓ Correct
- D.Training separate independent models for each individual language
Explanation
Cross-lingual transfer leverages knowledge learned from one language to improve performance on another, often using multilingual pre-trained models like mBERT or XLM-R.
Report an error in this question
NLPHard
Q24. What is the problem of hallucination in large language models?
- A.Generating plausible-sounding but factually incorrect or fabricated information✓ Correct
- B.Using too much GPU memory during the inference computation
- C.Producing completely empty outputs with no generated tokens
- D.Generating text output at an excessively slow processing speed
Explanation
Hallucination refers to LLMs generating confident, fluent text that contains fabricated facts, incorrect details, or unsupported claims.
Report an error in this question
NLPMedium
Q25. What is a language model?
- A.A static vocabulary dictionary for word lookup
- B.A rule-based grammar checker for syntax only
- C.A database of aligned translation memory pairs
- D.A model that predicts the probability of a sequence of words✓ Correct
Explanation
A language model estimates the probability distribution over sequences of words, used for text generation, completion, and evaluation of sentence fluency.
Report an error in this question
NLPHard
Q26. What is the concept of positional encoding in Transformers?
- A.Normalizing the positional coordinates of data in feature space
- B.Encoding the physical position of GPUs in a server rack cluster
- C.Adding position info to embeddings since Transformers lack inherent sequence order✓ Correct
- D.Sorting the input data samples by their index in the dataset
Explanation
Since Transformers process all positions in parallel (no recurrence), positional encodings add sequence order information to input embeddings using sinusoidal functions or learned embeddings.
Report an error in this question
NLPHard
Q27. What is retrieval-augmented generation (RAG)?
- A.Combining retrieval with generation to ground outputs in external knowledge✓ Correct
- B.An unsupervised clustering method for grouping similar documents
- C.A purely generative approach relying entirely on internal knowledge
- D.A purely retrieval approach returning documents without generation
Explanation
RAG retrieves relevant documents from a knowledge base and uses them as context for the language model's generation, reducing hallucination and enabling knowledge updates.
Report an error in this question
NLPHard
Q28. What is the difference between extractive and abstractive summarization?
- A.Extractive generates completely new text from the source material
- B.Abstractive only selects existing sentences from the original text
- C.Extractive and abstractive summarization are identical methods✓ Correct
- D.Extractive selects existing sentences; abstractive generates new sentences
Explanation
Extractive summarization selects key sentences from the source text. Abstractive summarization generates new sentences that capture the main ideas, potentially using words not in the original.
Report an error in this question
NLPHard
Q29. What is prompt engineering?
- A.Cleaning and preprocessing the raw input training data
- B.Modifying the internal neural network model architecture
- C.Training the language model from scratch on new domain data
- D.Designing effective input prompts to guide LLM behavior without fine-tuning✓ Correct
Explanation
Prompt engineering crafts input prompts that effectively guide pre-trained language models to produce desired outputs without modifying model weights.
Report an error in this question
NLPHard
Q30. What is the RLHF (Reinforcement Learning from Human Feedback) technique?
- A.Unsupervised pre-training using masked language modeling objectives
- B.Standard supervised fine-tuning using labeled input-output training pairs
- C.Data augmentation using paraphrasing and back-translation methods
- D.Fine-tuning language models using human preference judgments as reward signals✓ Correct
Explanation
RLHF trains a reward model from human preference comparisons and uses reinforcement learning to fine-tune the language model to generate outputs aligned with human preferences.
Report an error in this question