Python OS Get Files in Directory – Learn how to list, filter, and traverse files using the os module in Python with clear examples and performance tips.
Introduction
When working with file systems in Python, the os module provides a portable way to interact with directories and retrieve file names. Understanding python os get files in directory techniques is essential for tasks such as batch processing, data ingestion, or building utilities that need to enumerate contents of a folder. This guide walks you through the most common methods, shows how to apply filters, explains recursive traversal, and highlights best practices for reliable and efficient code.
Using os.listdir() to List Directory Contents
The simplest way to obtain all entries in a directory is with os.listdir(path). It returns a list containing the names of files and subdirectories (excluding the special . and .. entries).
import os
def list_all_entries(directory):
return os.listdir(directory)
# Example usage
entries = list_all_entries('/tmp/data')
print(entries)
Key points
- The function works with both relative and absolute paths.
- It returns only the names; you must join them with the directory path if you need full paths.
- The order of items is arbitrary and depends on the underlying file system.
Getting Full Paths with os.path.join()
Often you need the complete path to each item rather than just its name. Combining os.listdir() with os.path.join() achieves this:
import os
def get_full_paths(directory):
return [os.path.join(directory, name) for name in os.listdir(directory)]
full_paths = get_full_paths('/tmp/data')
print(full_paths)
Using os.path.join() ensures cross‑platform compatibility because it inserts the correct separator (/ on Unix, \ on Windows) That's the whole idea..
Filtering Files vs. Directories
Sometimes you want only files or only subdirectories. The os.path.isfile() and os.path.isdir() helpers let you test each entry:
import os
def get_files_only(directory):
return [name for name in os.isfile(os.Because of that, path. listdir(directory)
if os.path.
def get_dirs_only(directory):
return [name for name in os.listdir(directory)
if os.path.isdir(os.path.
files = get_files_only('/tmp/data')
dirs = get_dirs_only('/tmp/data')
print("Files:", files)
print("Directories:", dirs)
Tip: Wrap the logic in a function if you need to reuse it across multiple scripts.
Using os.scandir() for Better Performance
For large directories, os.scandir() is preferable because it yields DirEntry objects that expose file type information without extra system calls. This reduces overhead, especially when you need to filter by file type.
import os
def scan_files(directory):
with os.And scandir(directory) as it:
for entry in it:
if entry. is_file():
yield entry.
# Example
for file_path in scan_files('/tmp/data'):
print(file_path)
Advantages
DirEntryprovides attributes likename,path,inode, and methodsis_file(),is_dir(),is_symlink().- The iterator protocol means you can stop early without scanning the entire directory.
Recursive Directory Traversal
When you need to walk through a directory tree (i.e., include subdirectories), os.walk() is the standard tool. It generates a tuple (root, dirs, files) for each directory it visits It's one of those things that adds up. Turns out it matters..
import os
def list_all_files_recursive(start_dir):
file_list = []
for root, dirs, files in os.append(os.walk(start_dir):
for f in files:
file_list.path.
all_files = list_all_files_recursive('/tmp/data')
print(f"Found {len(all_files)} files")
Customizing the walk
- You can modify the
dirslist in‑place to prune branches you don’t want to visit (e.g., skip hidden folders). - To follow symlinks, pass
followlinks=Truetoos.walk().
def list_files_skip_hidden(start_dir):
file_list = []
for root, dirs, files in os.walk(start_dir, followlinks=False):
# Remove hidden directories from further traversal
dirs[:] = [d for d in dirs if not d.startswith('.')]
for f in files:
if not f.startswith('.'):
file_list.append(os.path.join(root, f))
return file_list
Error Handling and Edge Cases
File system operations can raise exceptions such as FileNotFoundError, PermissionError, or OSError. Wrapping calls in a try/except block makes your script dependable The details matter here..
import os
def safe_listdir(directory):
try:
return os.listdir(directory)
except FileNotFoundError:
print(f"Directory not found: {directory}")
return []
except PermissionError:
print(f"Permission denied: {directory}")
return []
except OSError as e:
print(f"OS error while accessing {directory}: {e}")
return []
contents = safe_listdir('/restricted/path')
When using os.scandir(), the context manager (with) automatically releases resources, but you still need to catch exceptions inside the block if you want custom handling Simple, but easy to overlook..
Performance Considerations
os.listdir()is fast for small to medium directories but requires extrastatcalls if you need to differentiate files from directories.os.scandir()reduces system calls by returningDirEntryobjects that already know the entry type, making it up to 2‑3× faster on large directories.os.walk()internally usesos.scandir()(since Python 3.5), so it inherits the same performance benefits.- Avoid repeatedly joining paths inside tight loops; pre‑compute the base path and use string concatenation or
os.path.join()only once per entry.
Best Practices Summary
- Choose the right tool –
os.listdir()for simple name lists,os.scandir()when you need type information,os.walk()for recursive walks. - Always join paths with
os.path.join()(orPathfrompathlibif you prefer an object‑oriented approach) to guarantee portability. - Filter early – apply file/directory tests as soon as possible to avoid unnecessary work.
- Handle errors gracefully – anticipate missing directories, permission issues, and other I/O problems.
- Release resources – use `scandir
Because os.scandir returns an iterator, it is advisable to wrap the call in a with statement; this guarantees that the underlying file descriptor is closed as soon as the block exits, preventing file‑descriptor leaks, especially when the directory contains a huge number of entries.
def list_files_with_scandir(root):
files = []
with os.scandir(root) as it:
for entry in it:
if not entry.name.startswith('.'):
if entry.is_file(follow_symlinks=False):
files.append(entry.path)
elif entry.is_dir(follow_symlinks=False):
files.extend(list_files_with_scandir(entry.path))
return files
The context manager ensures that the directory handle is released even if an exception occurs while iterating. If you need to process a directory tree without recursion, you can manually manage a stack of paths and invoke scandir for each entry, which gives you fine‑grained control over memory usage.
It sounds simple, but the gap is usually here.
When you prefer a more expressive API, pathlib.Consider this: it yields Path objects that already carry methods for common operations (e. Which means scandir. Path.g.iterdir()provides a thin wrapper aroundos., is_file(), is_dir()), allowing you to write code that reads naturally while still benefiting from the performance of the underlying iterator.
Additional considerations
- Permission errors – even with
scandir, aPermissionErrorcan be raised when trying to enter a subdirectory. Catching the exception inside the loop lets you skip problematic entries without aborting the whole traversal.
for entry in it:
try:
# …process entry…
except PermissionError:
continue # or log the issue and move on
-
Recursive depth – Python’s default recursion limit can be hit when walking very deep directory structures. Using
os.walk(topdown=False)or an explicit stack avoids this limitation Turns out it matters.. -
Pattern matching – for selective listing (e.g., only Python files), combine
scandirwithfnmatch.fnmatchorpathlib.Path.suffixchecks, applying the filter as early as possible to minimize work. -
Caching – if the same directory is accessed multiple times, consider storing the results of
scandirin a dictionary or usingfunctools.lru_cacheon a wrapper function that accepts a path string Less friction, more output..
Conclusion
Choosing the appropriate filesystem tool hinges on the task at hand: os.listdir is sufficient for quick name grabs, os.walk remains the most convenient option for recursive walks. scandirshines when entry types are needed and performance matters, andos.Wrapping iterators in context managers, handling permission and existence errors gracefully, and applying filters early in the processing pipeline all contribute to solid, maintainable code. By following these practices — selecting the right API, joining paths safely, filtering promptly, and releasing resources correctly — you can efficiently handle the file system while keeping your scripts reliable across diverse environments.