Project: Text Analyzer
Combine string methods, a loop, and a dictionary to summarize a sentence.
Video coming to YouTube
Work through the full lesson and run its code while the video is being prepared for publication.
Understand the concept
Start with text.lower().split(): lower makes case consistent, and split divides on whitespace. Count the resulting words with len. A dictionary can track words already seen, so its length gives the number of unique words.
For each word, increase its count with counts.get(word, 0) + 1. Return a small dictionary with total and unique counts. This project treats punctuation as part of a word: 'code' and 'code!' count separately. That simple rule keeps the first version predictable.
The function should also handle an empty string. split() returns an empty list, so both counts become zero. Once your checked solution works, try adding a third field for the most frequent word.
- lower() makes case-insensitive comparison possible
- split() separates whitespace-delimited words
- A dictionary records word frequencies
See it step by step
Read the code, predict the output, then compare it with the result.
01. Normalize and split
words = "Code code twice".lower().split()
print(words)['code', 'code', 'twice']lower turns both versions of Code into the same word; split creates a list.
02. Count total words
words = "one two three".split()
print(len(words))3The list has three items after splitting on spaces.
03. Count distinct words with a dictionary
counts = {}
for word in ["a", "b", "a"]:
counts[word] = counts.get(word, 0) + 1
print(counts)
print(len(counts)){'a': 2, 'b': 1}
2Dictionary keys are unique, so two a entries still create only one key.
A closer look
Follow the reasoning, inspect each result, then try the suggested changes in the console below.
Choose a definition of word
A word count depends on the rule you choose. Splitting on whitespace treats 'Python,' and 'Python' as different tokens because punctuation stays attached. A more deliberate rule can pull letter sequences out of the text first. The regular expression below keeps English letters and optionally an apostrophe within a word, so 'don't' remains one token. It deliberately does not handle every language or every kind of punctuation.
Normalize case after extracting tokens so different capitalization does not create extra dictionary keys. This is a product decision: for some reports, 'US' and 'us' should remain different. State the rule near the code and test it with examples that include punctuation, contractions, and an empty string. Try replacing the comma with an exclamation mark and predict whether the result changes.
import re
text = "Python, python! Don't stop."
tokens = re.findall(r"[A-Za-z]+(?:'[A-Za-z]+)?", text)
words = [token.lower() for token in tokens]
print(words)
print(len(words))['python', 'python', "don't", 'stop']
4- findall extracts matching token strings and ignores the comma, exclamation mark, and period.
- lower makes the two spellings of Python identical for counting.
- Test an empty string; the result should be an empty list and a count of zero.
Rank frequent words predictably
A frequency dictionary is useful when you want more than a total and a unique count. To show the most frequent words, sort pairs of (word, count) by descending count. Ties need a second rule; alphabetical order makes results stable even when the original sentence order changes. This matters for screenshots, tests, and readers who want the same example output each time.
Python's sorted function accepts a key that maps each pair to the values used for ordering. A negative count sorts large counts first, while the word itself breaks ties. The example returns a new list and leaves the counts dictionary available for other reports. Change the final 'red' to 'blue' and predict which word appears first.
words = "red blue red green blue red".split()
counts = {}
for word in words:
counts[word] = counts.get(word, 0) + 1
ranking = sorted(counts.items(), key=lambda pair: (-pair[1], pair[0]))
for word, count in ranking:
print(f"{word}: {count}")red: 3
blue: 2
green: 1- The dictionary stores one count per distinct word.
- The sorting key uses negative count for descending frequency and word for ties.
- Try input with equal counts, such as 'red blue', and predict alphabetical order.
Analyze several lines as one report
Text often arrives as more than one sentence. A line break is whitespace, so split() can count every word across lines, but line boundaries can still be useful. For example, you may want to know how many nonempty lines a note contains in addition to its word count. The code below records both without changing the original text.
Avoid calling strip() on the whole source if leading or trailing spacing is meaningful to your application. Here strip() is used only to decide whether a line has visible content. Splitting each line separately makes the decision easy to inspect. Try adding a line containing only spaces; the number of nonempty lines should stay the same.
text = "First line has words\n\nSecond line too\n"
nonempty_lines = 0
total_words = 0
for line in text.splitlines():
if line.strip():
nonempty_lines += 1
total_words += len(line.split())
print(f"Lines: {nonempty_lines}")
print(f"Words: {total_words}")Lines: 2
Words: 7- splitlines separates the source into three lines, including one blank line.
- The blank line is ignored; four words plus three words produce seven.
- Insert another blank or spaces-only line and predict both totals.
Try it in Python
Edit the example and run it. Python starts in your browser the first time you click Run.
Python console
Ready to runNeed input()? Add one value per line
Your output appears here.
Need a hint?
Make words = text.lower().split(). Start counts = {}. Loop over words and set counts[word] = counts.get(word, 0) + 1. Return {'words': len(words), 'unique': len(counts)}.
Complete the task and select Check task to verify your code.
Quick quiz
Three questions. You can change your answers and try again.
Typical mistakes
Everyone meets these errors. See what causes them and how to fix them.
Counting the text's characters
word_count = len(text)word_count = len(text.split())What happens: 'red blue' reports eight characters instead of two words.
Split the text first when you want a word count.
Treating capitalization as a new word
words = text.split()words = text.lower().split()What happens: 'Code' and 'code' become separate dictionary keys.
Normalize case before counting for a case-insensitive report.