Major Paradigms of Learning
Classical
Conditioning
Pavlov
Operant
Conditioning
Skinner
Observational
Learning
Bandura
Cognitive
Learning
Tolman, Köhler
Verbal &
Skill Learning
Language, motor
Fig 5.1 – Major paradigms of learning
Ivan Pavlov (1849–1936), a Russian physiologist, discovered classical conditioning while studying digestion in dogs.
Procedure:
Before conditioning: A bell (neutral stimulus) does not produce salivation. Food (unconditioned stimulus) naturally produces salivation.
During conditioning: The bell is repeatedly paired with food.
After conditioning: The bell alone (now a conditioned stimulus) produces salivation (conditioned response).
Classical Conditioning – Pavlov's Experiment
Before Conditioning
Bell (NS)
No salivation
Food (UCS)
Salivation (UCR)
During Conditioning
Bell (NS)
+
Food (UCS)
Salivation (UCR)
(repeated pairings)
After Conditioning
Bell (CS)
Salivation (CR)
Fig 5.2 – Classical conditioning process
Term Abbreviation Definition
Unconditioned Stimulus UCS Stimulus that naturally triggers a response (food)
Unconditioned Response UCR Natural response to the UCS (salivation to food)
Neutral Stimulus NS Stimulus that initially does not trigger the response (bell)
Conditioned Stimulus CS Previously neutral stimulus that, after pairing with UCS, triggers a response (bell after conditioning)
Conditioned Response CR Learned response to the CS (salivation to bell)
Factor Effect
Time relationship The CS should come slightly before the UCS (forward conditioning is most effective)
Frequency of pairing More pairings generally lead to stronger conditioning
Intensity of UCS A stronger UCS leads to faster conditioning
Salience of CS A more noticeable CS is conditioned more readily
Process Description
Acquisition The initial stage of learning (CS-UCS pairing)
Extinction If the CS is presented repeatedly without the UCS, the CR gradually weakens and disappears
Spontaneous recovery After extinction, the CR may briefly reappear when the CS is presented again
Stimulus generalisation The CR occurs in response to stimuli similar to the CS (e.g., a dog salivates to bells with different pitches)
Stimulus discrimination The organism learns to respond only to the specific CS, not to similar stimuli
Operant conditioning (also called instrumental conditioning) is learning based on the consequences of behaviour. It was developed by B.F. Skinner (1904–1990).
Core principle: Behaviours followed by positive consequences (reinforcement) tend to be repeated; behaviours followed by negative consequences (punishment) tend to decrease.
Type Definition Example
Positive reinforcement Adding something pleasant after a behaviour → behaviour increases A student gets praise for answering correctly → answers more
Negative reinforcement Removing something unpleasant after a behaviour → behaviour increases Taking an aspirin removes a headache → more likely to take aspirin next time
Positive punishment Adding something unpleasant after a behaviour → behaviour decreases A child touches a hot stove → gets burned → stops touching
Negative punishment Removing something pleasant after a behaviour → behaviour decreases A teenager breaks curfew → loses phone privileges → follows curfew
Reinforcement vs. Punishment
Add Stimulus (+)
Remove Stimulus (−)
Behaviour ↑
(Increases)
Behaviour ↓
(Decreases)
Positive
Reinforcement
Negative
Reinforcement
Positive
Punishment
Negative
Punishment
Fig 5.3 – Grid of reinforcement and punishment types
Factor Description
Immediacy Reinforcement/punishment is more effective when delivered immediately after the behaviour
Consistency Consistent reinforcement establishes behaviour faster
Schedule of reinforcement Continuous vs. partial (intermittent) reinforcement affects persistence
Schedule Description Example
Continuous Reinforcement after every correct response Giving a treat every time a dog sits on command
Fixed ratio Reinforcement after a fixed number of responses Getting paid for every 10 items produced
Variable ratio Reinforcement after an unpredictable number of responses Slot machines (most resistant to extinction)
Fixed interval Reinforcement for the first response after a fixed time period Monthly salary
Variable interval Reinforcement for the first response after an unpredictable time period Surprise quizzes
Process Description
Shaping Reinforcing successive approximations to the desired behaviour
Chaining Linking a series of individual behaviours to form a complex sequence
Extinction Withholding reinforcement leads to gradual decrease of the behaviour
Generalisation A behaviour reinforced in one situation occurs in similar situations
Discrimination Learning to respond only in specific situations where reinforcement is available
Albert Bandura proposed that people learn by watching others (models) — without direct reinforcement.
Children who observed an adult behaving aggressively towards a Bobo doll later imitated the aggressive behaviour, especially when the adult was rewarded for the aggression.
Attention — The learner must pay attention to the model’s behaviour
Retention — The learner must remember what was observed
Reproduction — The learner must be capable of reproducing the behaviour
Motivation — The learner must have a reason to imitate (e.g., expected reward)
Vicarious reinforcement: Learning that observed behaviour leads to positive consequences increases the likelihood of imitation.
Cognitive approaches emphasise the role of mental processes (thinking, expectation, insight) in learning.
Köhler studied chimpanzees who had to reach bananas placed out of reach
The chimps suddenly figured out how to stack boxes or join sticks to reach the bananas — an “aha!” moment
Insight involves sudden understanding of a problem’s solution, not gradual trial-and-error
Latent learning is learning that occurs without any obvious reinforcement and is not demonstrated until there is an incentive to do so
Tolman’s maze experiment: Rats that explored a maze without reward later learned the maze faster when food was introduced — they had formed a cognitive map (mental representation of the maze)
Verbal learning involves learning through language — words, sentences, and texts.
Methods used to study verbal learning:
Method Description
Paired associate learning Learning to associate pairs of words (e.g., DOG – BLUE)
Serial learning Learning a list of items in a specific order
Free recall Learning a list and recalling items in any order
Skill learning involves acquiring coordinated motor or cognitive patterns through practice.
Phase Description
Cognitive phase Understanding what needs to be done (learning rules and procedures)
Associative phase Practising and refining the skill; errors decrease
Autonomous phase The skill becomes automatic; requires minimal conscious effort
Example: Learning to ride a bicycle — first you learn the concept (pedal, steer, balance), then you practise, and eventually it becomes automatic.
Factor Effect
Motivation Higher motivation leads to better learning
Reinforcement Behaviours followed by positive outcomes are learned faster
Practice Repeated practice strengthens learning (distributed practice is better than massed practice)
Meaningfulness Meaningful material is learned more easily than meaningless material
Prior knowledge Existing knowledge provides a framework for new learning
Feedback Knowledge of results helps learners correct errors
Active participation Learning by doing is more effective than passive reading
Organisation Organised material is easier to learn and recall
Learning disabilities are neurologically based processing problems that can interfere with learning basic skills such as reading, writing, or mathematics.
Disability Description
Dyslexia Difficulty with reading — problems with word recognition, spelling, decoding
Dysgraphia Difficulty with writing — poor handwriting, spelling, organising ideas on paper
Dyscalculia Difficulty with mathematics — problems with number sense, calculations, mathematical reasoning
ADHD (Attention Deficit Hyperactivity Disorder)Difficulty sustaining attention, hyperactivity, impulsivity — affects learning
Learning disabilities are not a sign of low intelligence. Individuals with learning disabilities often have average or above-average intelligence but process information differently.
Term Meaning
Learning Relatively permanent change in behaviour due to experience
Classical conditioning Learning through association of stimuli (Pavlov)
Operant conditioning Learning based on consequences of behaviour (Skinner)
Reinforcement Consequence that increases the probability of a behaviour
Punishment Consequence that decreases the probability of a behaviour
Observational learning Learning by watching others (Bandura)
Insight Sudden understanding of a problem’s solution (Köhler)
Latent learning Learning without reinforcement, shown later (Tolman)
Shaping Reinforcing successive approximations
Extinction Weakening of a response when reinforcement is withheld
Q1. In Pavlov’s experiment, the bell after conditioning is called the:
(a) Unconditioned stimulus (b) Conditioned stimulus (c) Neutral stimulus (d) Unconditioned response
Answer
(b) Conditioned stimulus (CS)
After repeated pairing with the unconditioned stimulus (food), the previously neutral bell becomes a conditioned stimulus that can trigger salivation on its own.
Q2. A child stops misbehaving in class after the teacher takes away recess privileges. This is an example of:
(a) Positive reinforcement (b) Negative reinforcement (c) Positive punishment (d) Negative punishment
Answer
(d) Negative punishment
Something pleasant (recess) is removed after the behaviour, causing the behaviour (misbehaving) to decrease . This is negative punishment.
Q3. Bandura’s Bobo doll experiment demonstrated:
(a) Classical conditioning (b) Operant conditioning (c) Observational learning (d) Insight learning
Answer
(c) Observational learning
Children learned aggressive behaviour by watching an adult model, without being directly reinforced themselves. This demonstrated Bandura’s theory of observational (social) learning.
Q4. A slot machine that pays out after an unpredictable number of plays uses a:
(a) Fixed ratio schedule (b) Variable ratio schedule (c) Fixed interval schedule (d) Variable interval schedule
Answer
(b) Variable ratio schedule
Slot machines reinforce after an unpredictable number of responses. Variable ratio schedules produce the highest and most consistent rate of responding and are most resistant to extinction .
Q5. Which of the following is NOT a learning disability?
(a) Dyslexia (b) Dyscalculia (c) Dysgraphia (d) Dyspepsia
Answer
(d) Dyspepsia
Dyspepsia is a medical condition (indigestion), not a learning disability. Dyslexia (reading), dyscalculia (mathematics), and dysgraphia (writing) are all learning disabilities.
Q6. Distinguish between classical conditioning and operant conditioning.
Answer
Feature Classical Conditioning Operant Conditioning
Pioneer Ivan Pavlov B.F. Skinner
Nature Learning through association of two stimuli Learning through consequences of behaviour
Response Involuntary / reflexive (e.g., salivation) Voluntary (e.g., pressing a lever)
Role of learner Passive — responds to stimuli Active — operates on the environment
Reinforcement UCS serves as reinforcement; precedes the response Follows the response; determines whether it increases or decreases
Q7. Explain the concepts of reinforcement and punishment with examples.
Answer
Reinforcement is any consequence that increases the frequency of a behaviour.
Positive reinforcement: Adding something pleasant. E.g., A teacher gives a star to a student who completes homework → the student completes homework more regularly.
Negative reinforcement: Removing something unpleasant. E.g., A seat belt alarm stops buzzing when you buckle up → you buckle up more quickly.
Punishment is any consequence that decreases the frequency of a behaviour.
Positive punishment: Adding something unpleasant. E.g., A child is scolded for lying → the child lies less.
Negative punishment: Removing something pleasant. E.g., A teenager loses screen time for poor grades → the teenager studies more.
Q8. What is latent learning? Describe Tolman’s experiment.
Answer
Latent learning is learning that occurs without any obvious reinforcement and is not demonstrated until there is a motivation or incentive to display it.
Tolman’s maze experiment:
Three groups of rats were placed in a maze over several days
Group 1: Rewarded with food every day — learned the maze steadily
Group 2: Never rewarded — showed little improvement
Group 3: Not rewarded for the first 10 days, then rewarded from day 11
On day 11, Group 3 suddenly performed as well as Group 1, demonstrating that they had already learned the maze layout (formed a cognitive map ) during the unrewarded period. They simply had no motivation to demonstrate it earlier.
This showed that learning can occur without reinforcement and that mental representations play a role in learning.
Q9. Describe the process of classical conditioning with reference to Pavlov’s experiment. Discuss the key processes involved.
Answer
Classical conditioning is a type of learning in which a previously neutral stimulus comes to elicit a response after being paired with an unconditioned stimulus.
Pavlov’s Experiment:
Before conditioning: Food (UCS) naturally produces salivation (UCR). A bell (NS) produces no salivation.
During conditioning: The bell is sounded just before food is presented. This pairing is repeated multiple times.
After conditioning: The bell alone (now CS) produces salivation (now CR).
Key Processes:
Acquisition: The stage during which the CS-UCS association is established through repeated pairings. The strength of the CR increases with each pairing.
Extinction: If the CS (bell) is presented repeatedly without the UCS (food), the CR (salivation) gradually weakens and eventually disappears.
Spontaneous recovery: After a rest period following extinction, the CS may again produce the CR, though it is usually weaker. This shows that extinction does not completely erase the learned association.
Stimulus generalisation: The CR may be triggered by stimuli similar to the CS. E.g., Pavlov’s dog may salivate to bells of different pitches.
Stimulus discrimination: With training, the organism learns to distinguish between the CS and similar stimuli, responding only to the specific CS.
Real-life example: A child bitten by a dog (UCS) develops fear (UCR). Later, the child fears all dogs (stimulus generalisation through classical conditioning).
Q10. Explain observational learning with reference to Bandura’s theory and experiment.
Answer
Observational learning (also called social learning or modelling) is the process of learning by watching the behaviour of others (models) and noting the consequences of that behaviour.
Albert Bandura’s Social Learning Theory proposes that a significant amount of human learning occurs through observing others, without direct personal reinforcement.
Bobo Doll Experiment (1961):
Children were divided into groups and shown an adult model interacting with a Bobo doll (inflatable toy)
Group 1: Observed an adult behaving aggressively (punching, kicking, hitting the doll with a mallet)
Group 2: Observed an adult playing calmly
Group 3: No model (control group)
When later left alone with the Bobo doll, children who observed the aggressive model displayed significantly more aggressive behaviour
Four Steps in Observational Learning:
Attention: The learner must attend to the model’s behaviour. Factors: attractiveness of model, similarity, status
Retention: The learner must remember (encode) the observed behaviour in memory
Motor reproduction: The learner must be physically and cognitively capable of performing the behaviour
Motivation: The learner must be motivated to imitate. Motivation is enhanced by:
Vicarious reinforcement — seeing the model being rewarded
Self-reinforcement — personal satisfaction from performing the behaviour
Vicarious punishment — seeing the model punished decreases imitation
Significance:
Explains how children learn social behaviours, gender roles, and aggression
Highlights the influence of media, parents, and peers as models
Has applications in education, therapy (modelling adaptive behaviours), and understanding the effects of media violence
Q11. (Assertion–Reason)
Assertion (A): Negative reinforcement increases the probability of a behaviour.
Reason (R): Negative reinforcement involves removal of an unpleasant stimulus after the desired behaviour.
(a) Both A and R are true and R is the correct explanation of A
(b) Both A and R are true but R is NOT the correct explanation of A
(c) A is true but R is false
(d) A is false but R is true
Answer
(a) Both A and R are true and R is the correct explanation of A
Negative reinforcement increases behaviour by removing an aversive (unpleasant) stimulus after the behaviour occurs. The removal of discomfort makes the behaviour more likely to be repeated. For example, taking a painkiller (behaviour) removes a headache (unpleasant stimulus), so you are more likely to take it again in the future.
Q12. (Case Study)
Arjun is a 7-year-old boy. His parents want him to clean his room regularly. They try the following strategies over three weeks:
Week 1: They give Arjun a star sticker every time he cleans his room.
Week 2: They scold him whenever he does not clean his room.
Week 3: They stop giving stickers but do not scold him either. Arjun gradually stops cleaning his room.
(i) Identify the type of operant conditioning used in Week 1 and Week 2.
(ii) What process is illustrated in Week 3?
(iii) Suggest a more effective long-term strategy using operant conditioning principles.
(iv) Why is Week 1’s approach generally more effective than Week 2’s?
Answer
(i)
Week 1: Positive reinforcement — a pleasant stimulus (star sticker) is added after the desired behaviour (cleaning the room), increasing its frequency.
Week 2: Positive punishment — an aversive stimulus (scolding) is added after the undesired behaviour (not cleaning), aiming to decrease it.
(ii) Week 3 illustrates extinction — the reinforcement (stickers) is no longer provided, so the previously reinforced behaviour (cleaning) gradually decreases and eventually stops.
(iii) A more effective long-term strategy would be to use a variable ratio schedule of reinforcement :
Instead of rewarding every single instance, give rewards unpredictably (e.g., stickers sometimes, verbal praise other times)
This creates a pattern that is highly resistant to extinction
Also, gradually shifting from external reward to intrinsic motivation (e.g., praising the feeling of having a clean room) would help sustain the behaviour
(iv) Week 1 (positive reinforcement) is generally more effective because:
It teaches the child what to do (clean the room), not just what not to do
It creates a positive association with the behaviour
Punishment (Week 2) can create fear, resentment, and avoidance without teaching alternative behaviour
Research shows reinforcement produces more lasting behavioural change than punishment
Q13. (Application-Based)
A teacher wants to teach a pigeon to turn in a complete circle. The pigeon has never performed this behaviour before.
(i) Which operant conditioning technique should the teacher use?
(ii) Describe the step-by-step procedure the teacher would follow.
(iii) What type of reinforcement schedule should be used initially?
Answer
(i) The teacher should use shaping — the method of reinforcing successive approximations to the desired behaviour.
(ii) Step-by-step procedure:
First, reinforce the pigeon for any slight turn of the head in the desired direction
Once head-turning is established, withhold reinforcement until the pigeon turns its body slightly
Next, reinforce only a quarter turn
Then reinforce only a half turn
Then reinforce a three-quarter turn
Finally, reinforce only a complete 360-degree turn
At each step, the criterion for reinforcement becomes closer to the final desired behaviour.
(iii) Initially, a continuous reinforcement schedule (reinforcing every correct response) should be used. This helps the pigeon learn the new behaviour quickly. Once the behaviour is established, the teacher can gradually shift to a partial (intermittent) reinforcement schedule to make the behaviour more resistant to extinction.