Android Race Condition Exploit - CVE-2022-22057
Root Cause Analysis
The kgsl_timeline object can be created and destroyed through the IOCTL_KGSL_TIMELINE_CREATE / IOCTL_KGSL_TIMELINE_DESTROY ioctls.
struct kgsl_timeline {
/** @context: dma-fence timeline context */
u64 context;
/** @id: Timeline identifier */
int id;
/** @value: Current value of the timeline */
u64 value;
/** @fence_lock: Lock to protect @fences */
spinlock_t fence_lock;
/** @lock: Lock to use for locking each fence in @fences */
spinlock_t lock;
/** @ref: Reference count for the struct */
struct kref ref;
/** @fences: sorted list of active fences */
struct list_head fences;
/** @name: Name of the timeline for debugging */
const char name[32];
/** @dev_priv: pointer to the owning device instance */
struct kgsl_device_private *dev_priv;
};
The kgsl_timeline object stores a list of dma_fence objects in its fences field.
The IOCTL_KGSL_TIMELINE_FENCE_GET / IOCTL_KGSL_TIMELINE_WAIT ioctls can add dma_fence objects to that list.
The added dma_fence objects are refcounted objects, and their refcount is decremented via the dma_fence_put function.
What's interesting is that, separate from the fact that dma_fence has its own refcount, the timeline→fences list where they are stored does not have a separate refcount.
Instead, to prevent a dma_fence in the timeline→fences field from being freed, a dedicated release function called timeline_fence_release was created, which removes the dma_fence from timeline→fences before the dma_fence is freed.
static void timeline_fence_release(struct dma_fence *fence)
{
...
spin_lock_irqsave(&timeline->fence_lock, flags);
/* If the fence is still on the active list, remove it */
list_for_each_entry_safe(cur, temp, &timeline->fences, node) {
if (f != cur)
continue;
list_del_init(&f->node); //<----- 1. Remove fence
break;
}
spin_unlock_irqrestore(&timeline->fence_lock, flags);
...
kgsl_timeline_put(f->timeline);
dma_fence_free(fence); //<------- 2. frees the fence
}
When the refcount of a dma_fence stored in kgsl_timeline→fences reaches 0, the kref_put function calls the timeline_fence_release function to remove the dma_fence object from the fences list.
After the dma_fence object is removed, it is freed via dma_fence_free.
Additionally, a spinlock protects the removal from the fences list against race conditions.
However, through the IOCTL_KGSL_TIMELINE_DESTROY ioctl, it is possible to acquire a reference to a dma_fence after its refcount has reached 0 but before it has been removed in timeline_fence_release.
long kgsl_ioctl_timeline_destroy(struct kgsl_device_private *dev_priv,
unsigned int cmd, void *data)
{
...
// critical section 1
spin_lock(&timeline->fence_lock);
list_for_each_entry_safe(fence, tmp, &timeline->fences, node)
dma_fence_get(&fence->base);
list_replace_init(&timeline->fences, &temp);
spin_unlock(&timeline->fence_lock);
// critical section 2
spin_lock_irq(&timeline->lock);
list_for_each_entry_safe(fence, tmp, &temp, node) {
dma_fence_set_error(&fence->base, -ENOENT);
dma_fence_signal_locked(&fence->base);
dma_fence_put(&fence->base);
}
spin_unlock_irq(&timeline->lock);
kgsl_timeline_put(timeline);
return timeline ? 0 : -ENODEV;
}
The kgsl_ioctl_timeline_destroy function goes through the following steps to destroy the timeline object.
- As described earlier, since the
fences list does not have a separate refcount, it iterates over fences and increments the refcount of every dma_fence object inside by 1 in preparation for the subsequent work.
- It copies the
fences list into a new list called temp and reinitializes fences.
- It iterates again, this time decrementing the refcount of the
dma_fence objects in temp by 1.
- Finally, it decrements the refcount of the
timeline object by 1 and finishes.
This logic, too, protects its operations from race conditions using a spinlock.
However, if it reaches step 2 while the refcount has become 0 but the timeline_fence_release function has not yet removed the dma_fence, then the fences list will be copied into temp.
Furthermore, even though the kgsl_ioctl_timeline_destroy function incremented the refcount prior to the copy, the already-invoked timeline_fence_release function belatedly removes the dma_fence and then frees it, so a UAF occurs.
The figure above illustrates the race condition situation.
The red blocks are mutually exclusive blocks — that is, blocks that hold the same lock.
Therefore, Thread 2's timeline_fence_release function cannot execute its red block — that is, the removal and freeing of the dma_fence — until Thread 1's kgsl_ioctl_timeline_destroy function releases the lock after copying the fences field into temp in its own red block.
By increasing the number of dma_fence objects in the fences list, we can lengthen the time it takes for Thread 1's kgsl_ioctl_timeline_destroy function to iterate over the fences list and increment the refcount of all dma_fence objects.
If, while Thread 1's red block is executing, Thread 2 decrements the refcount of the last dma_fence object in the fences list to 0, then before Thread 1's iteration reaches the last dma_fence object and increments its refcount, that object's refcount will become 0, and the kref_put function will invoke the timeline_fence_release function as a callback.
Thread 2's red block corresponding to the code of the timeline_fence_release function must acquire the timeline→fence_lock that is currently held by Thread 1's red block, so it cannot be invoked until Thread 1's red block has finished executing.
Meanwhile, Thread 1's dma_fence objects in the fences list are moved into the temp list and the fences list is reinitialized. Therefore, when Thread 2's red block executes, the fences list is empty, so it finishes the iteration quickly and calls the dma_fence_free function.
long kgsl_ioctl_timeline_destroy(struct kgsl_device_private *dev_priv,
unsigned int cmd, void *data)
{
...
spin_lock_irq(&timeline->lock);
list_for_each_entry_safe(fence, tmp, &temp, node) {
dma_fence_set_error(&fence->base, -ENOENT);
dma_fence_signal_locked(&fence->base);
dma_fence_put(&fence->base);
}
spin_unlock_irq(&timeline->lock);
kgsl_timeline_put(timeline);
return timeline ? 0 : -ENODEV;
}
Afterwards, Thread 1 re-acquires the released lock and iterates over the fences list that was moved into temp, decrementing the refcount of the dma_fence objects in the list by 1.
At this point, the very last object in the list has already been freed by Thread 2, so we can reallocate an arbitrary object and manipulate its data. Thus, if we set that object's refcount to 1, then upon the refcount decrement the refcount becomes 0, and it will attempt to free it again via the timeline_fence_release function.
Therefore, simply by increasing the number of dma_fence objects in the fences list, we can trigger the race condition without much difficulty, and as long as we can decrement the refcount of the last dma_fence object in the fences list, we can trigger a UAF — and, going further, a DFB (Double Free Bug).
Proof of Concept
Adding dma_fence
First, we need to add enough dma_fence objects to the kgsl_timeline->fences list to make the race condition occur.
There are two ways to do this.
IOCTL_KGSL_TIMELINE_FENCE_GET
long kgsl_ioctl_timeline_fence_get(struct kgsl_device_private *dev_priv,
unsigned int cmd, void *data)
{
...
timeline = kgsl_timeline_by_id(device, param->timeline);
...
fence = kgsl_timeline_fence_alloc(timeline, param->seqno); //<----- dma_fence created and added to timeline
...
sync_file = sync_file_create(fence);
if (sync_file) {
fd_install(fd, sync_file->file);
param->handle = fd;
} else {
put_unused_fd(fd);
ret = -ENOMEM;
}
out:
dma_fence_put(fence);
kgsl_timeline_put(timeline);
return ret;
}
The function above calls kgsl_timeline_fence_alloc to create and initialize a fence, and then within that function calls kgsl_timeline_add_fence again to link the created fence object to the kgsl_timeline object.
After that, it obtains a sync_file fd for the dma_fence object via the sync_file_create function.
Also, by closing that fd, we can decrement the refcount of the dma_fence.
IOCTL_KGSL_TIMELINE_WAIT
long kgsl_ioctl_timeline_wait(struct kgsl_device_private *dev_priv,
unsigned int cmd, void *data)
{
...
fence = kgsl_timelines_to_fence_array(device, param->timelines,
param->count, param->timelines_size,
(param->flags == KGSL_TIMELINE_WAIT_ANY)); //<------ dma_fence created and added to timeline
...
if (!timeout)
ret = dma_fence_is_signaled(fence) ? 0 : -EBUSY;
else {
ret = dma_fence_wait_timeout(fence, true, timeout); //<----- 1.
...
}
dma_fence_put(fence);
...
}
The function above calls the kgsl_timelines_to_fence_array function, which, like the previous function, internally calls kgsl_timeline_fence_alloc again and links the fence object to the kgsl_timeline object through the same process as the previous function.
If the timeout value is not 0, this function executes the dma_fence_wait_timeout function and waits until the timeout expires or an interrupt occurs.
Therefore, if we set the timeout value very large, it will free the dma_fence object linked to the timeline when an interrupt occurs.
The kgsl_ioctl_timeline_fence_get function is convenient to use, but freeing a dma_fence in that function requires closing the sync_file fd, which incurs a lot of overhead.
Therefore, the last dma_fence that needs to be freed will be allocated via the kgsl_ioctl_timeline_wait function, while the other, non-freed dma_fence objects that fill the slots for the race condition will be allocated via the kgsl_ioctl_timeline_fence_get function.
Race Condition
To restate the analysis from the Root Cause Analysis section: while Thread 1's red block is executing, we must decrement the refcount of the last dma_fence object in the fences list to 0 so that the timeline_fence_release function is invoked.
Also, as mentioned earlier, by linking enough dma_fence objects to the fences list, we can lengthen the execution time of the red block.
However, Thread 2's red block must also complete before all of the remaining code of Thread 1 executes.
One might argue that because of the spinlock, the rest of Thread 1's code won't execute until Thread 2's red block finishes — but the most important dma_fence_free function is not protected by the spinlock, and in Thread 2's iteration the fences list is empty, so the code finishes executing very quickly. Crucially, dma_fence_free uses the kfree_rcu function, so the actual freeing of the object is deferred by the scheduler.
For this reason, unless we manipulate the scheduler, freeing the last dma_fence object within an appropriate timeframe is nearly impossible.
To solve this, we will use the technique introduced in LSSEU2019 - Exploiting race conditions on [ancient] Linux.
The content is as follows.
Race Condition in Tiny Race window
To guarantee each task sufficient CPU occupancy time, the Linux kernel can raise an interrupt during work and put a task to wait so that another task occupies the CPU.
This is called preemption. A task can also wait on its own and let another task preempt the CPU (e.g., waiting for I/O input, calling the sched_yield function).
Such preemption can occur even inside a syscall, such as an ioctl call.
And in the Android kernel, such preemption can occur except in a few special situations (e.g., while holding a spinlock).
This behavior can be manipulated via CPU affinity and task priority.
In the normal case, a task runs with SCHED_NORMAL priority.
However, the lower SCHED_IDLE priority can also be applied via the sched_setscheduler function (or, for threads, the pthread_setschedparam function).
Also, via the sched_setaffinity function, we can specify the CPU on which the code will run.
Therefore, by pinning two tasks to a single CPU and setting one to SCHED_NORMAL priority and the other to SCHED_IDLE priority, it is possible to manipulate preemption timing in the following way.
- The
SCHED_NORMAL task calls a syscall that would yield the CPU. For example, if it reads an empty pipe, it waits until data arrives and yields the CPU. Consequently, the SCHED_IDLE task that received the yield occupies the CPU.
- The
SCHED_IDLE task sends data to the pipe on which the SCHED_NORMAL task is waiting for data. As a result, the higher-priority SCHED_NORMAL wakes up from its wait state and preempts the CPU occupied by the SCHED_IDLE task. Consequently, the SCHED_IDLE task waits again.
SCHED_NORMAL continues its work so that the SCHED_IDLE task does not get the CPU back.
In the case of the PoC code, it is as follows.
- Call the
kgsl_ioctl_timeline_wait function on one thread to add a dma_fence object to the kgsl_timeline object. Set the timeout large and, via the sched_setaffinity function, pin this work to a single CPU that we'll name SPRAY_CPU.
As analyzed in IOCTL_KGSL_TIMELINE_WAIT, if a timeout is set, it waits until the timeout expires or an interrupt occurs after the dma_fence object is added.
- Create a
SCHED_NORMAL task and pin this work to another CPU that we'll name DESTROY_CPU.
This task reads an empty pipe to enter a wait state and yield the CPU to the lower-priority task.
Later, when data arrives in the empty pipe, this task continues its work so that the CPU is not yielded.
- Create a
SCHED_IDLE task and pin it to DESTROY_CPU.
This task calls the kgsl_ioctl_timeline_destroy function to destroy the kgsl_timeline object containing the dma_fence added in step 1. Since the task from step 2 is in a wait state, DESTROY_CPU runs this work first.
- On another CPU, when the
kgsl_ioctl_timeline_destroy function is iterating over the fences list in the aforementioned red block, raise an interrupt on the task running the kgsl_ioctl_timeline_wait function to end that task's wait state.
- Having received the interrupt, the
kgsl_ioctl_timeline_wait function decrements the refcount of the dma_fence it allocated. Because the kgsl_ioctl_timeline_destroy function, while iterating over the fences list, has not yet incremented the refcount of the dma_fence object added by kgsl_ioctl_timeline_wait, that dma_fence object's refcount becomes 0 and the timeline_fence_release function is invoked. Of course, since the spinlock is still held by kgsl_ioctl_timeline_destroy, it cannot perform its actual operation yet.
- The
SCHED_NORMAL task writes data to the empty pipe on which it is waiting. As a result, the SCHED_NORMAL task that had been waiting for the pipe's data escapes its wait state and preempts the CPU occupied by SCHED_IDLE. After preempting, it continues its work so that SCHED_IDLE does not get the CPU yielded back to it.
- In the
SCHED_IDLE task, the kgsl_ioctl_timeline_destroy function, which had been holding the spinlock, finishes copying and reinitializing the fences list into the temp list, and as soon as it releases the spinlock, it is preempted by the SCHED_NORMAL task and waits.
- Since the
kgsl_ioctl_timeline_destroy function released the spinlock, the timeline_fence_release function, which had been waiting for that spinlock, acquires the spinlock, quickly finishes iterating over the empty fences list, and calls the dma_fence_free function, which attempts to free — via the kfree_rcu function — the dma_fence object that kgsl_ioctl_timeline_wait had allocated. Of course, freeing via kfree_rcu has a delay until the actual free, but since we can control the duration of kgsl_ioctl_timeline_destroy's wait state, this delay is not a problem.
- After the
kfree_rcu free completes, the SCHED_NORMAL task again yields the CPU to the SCHED_IDLE task so that the remaining code of the kgsl_ioctl_timeline_destroy function executes.
Object Replacement
For the freed dma_fence object, we will use sendmsg, which is frequently used for heap spraying. It doesn't matter much, as long as we can write arbitrary data to the needed portion of whatever object we use.
Now let's figure out how to manipulate the reallocated object.
long kgsl_ioctl_timeline_destroy(struct kgsl_device_private *dev_priv,
unsigned int cmd, void *data)
{
...
spin_lock_irq(&timeline->lock);
list_for_each_entry_safe(fence, tmp, &temp, node) {
dma_fence_set_error(&fence->base, -ENOENT);
dma_fence_signal_locked(&fence->base);
dma_fence_put(&fence->base);
}
spin_unlock_irq(&timeline->lock);
kgsl_timeline_put(timeline);
return timeline ? 0 : -ENODEV;
}
The code above is the remaining code of the kgsl_ioctl_timeline_destroy function mentioned in step 9.
This code calls the dma_fence_set_error, dma_fence_signal_locked, and dma_fence_put functions with the now-reallocated object fence as the argument.
The dma_fence_set_error function writes an error code into the fence object. We could overwrite a specific field of a specific object with an error code and use it for exploitation, but in this PoC we won't explore that possibility.
Next, the code of the dma_fence_signal_locked function is as follows.
int dma_fence_signal_locked(struct dma_fence *fence)
{
...
if (unlikely(test_and_set_bit(DMA_FENCE_FLAG_SIGNALED_BIT, //<-- 1.
&fence->flags)))
return -EINVAL;
/* Stash the cb_list before replacing it with the timestamp */
list_replace(&fence->cb_list, &cb_list); //<-- 2.
...
list_for_each_entry_safe(cur, tmp, &cb_list, node) { //<-- 3.
INIT_LIST_HEAD(&cur->node);
cur->func(fence, cur);
}
return 0;
}
This function first checks: if DMA_FENCE_FLAG_SIGNALED_BIT is set in fence→flags, then the fence has already been signaled, so the function terminates early. (marked as 1 in the comments)
If not, the list_replace function is called to remove the objects in fence→cb_list and store them in cb_list. (marked as 2 in the comments)
After that, the function pointers of the objects in cb_list are called. (marked as 3 in the comments)
It would be nice if we could use this to call an arbitrary function, but since kCFI is applied, we cannot call anything that isn't a related function. Also, since we don't currently know function addresses, not only can we not call an arbitrary function, we also can't manipulate the reallocated object to make it call a legitimate function so as to avoid a crash.
Therefore, we have no choice but to set DMA_FENCE_FLAG_SIGNALED_BIT in fence→flags to terminate the function quickly.
The last function is the dma_fence_put function. As we saw earlier, this function calls the dma_fence_release function when the refcount of the fence object reaches 0.
void dma_fence_release(struct kref *kref)
{
...
if (fence->ops->release)
fence->ops->release(fence);
else
dma_fence_free(fence);
}
When the dma_fence_release function is called, at some point it checks fence→ops and calls the fence→ops→release function. This causes two problems.
The first problem is that fence→ops must point to valid memory. Otherwise, the dereference fails.
The second problem is that even if the dereference succeeds, fence→ops→release must be either 0 or the address of an appropriate function that doesn't trip kCFI.
There are two options for solving these problems.
Follow the standard approach and use a different object instead of the sendmsg object, or make use of the limited write primitive provided by the dma_fence_put and dma_fence_set_error functions to try to manipulate the flags and refcount fields so as to prevent the kernel from crashing due to the dma_fence_signal_locked or dma_fence_release functions.
Or we could try something else.
On Android, the ion allocator is used to allocate memory regions for DMA, which allow the kernel driver and a user process to share the same memory. This ion allocator is accessible from an untrusted app via the /dev/ion file, and an ion buffer can be allocated via the ION_IOC_ALLOC ioctl.
This ioctl returns an fd to the user, and this fd can be used via the mmap function to map the ion buffer's backing store into userspace.
What makes this ion buffer special is that the user can request Kernel Low Memory.
Kernel memory is divided into Kernel Low Memory and Kernel High Memory.
Here, unlike Kernel High Memory, Kernel Low Memory has a one-to-one mapping to physical addresses.
Therefore, the virtual address of that region is the physical address plus a fixed offset.
To obtain contiguous physical addresses, the ion driver allocates this memory region early during boot and uses it as a memory pool. This memory pool is later used to allocate ion buffers upon request.
Of course, not all memory pools of the ion device are contiguous, but when using the ION_IOC_ALLOC ioctl, we can request contiguous physical addresses by passing a specific heap_id_mask.
Because of these characteristics, if we request an ion buffer allocation from a rarely-used memory pool, we can predict that buffer's address, and by mapping the buffer into userspace via mmap, we can access a predictable address at any time.
According to prior research experiments, the user_contig_region region is almost never used, and it was possible to map the entire region into userspace.
Now that we have obtained both a controllable memory region and that region's address, we can solve the problem.
If we make fence→ops point to an ion buffer at a predictable address and zero-initialize that buffer,
then all of the aforementioned conditions are satisfied, the dma_fence_free function is called, and we can obtain an additional DFB primitive for fence.
But there is one more problem to solve.
Escaping an infinite loop
spin_lock_irq(&timeline->lock);
list_for_each_entry_safe(fence, tmp, &temp, node) {
dma_fence_set_error(&fence->base, -ENOENT);
dma_fence_signal_locked(&fence->base);
dma_fence_put(&fence->base);
}
spin_unlock_irq(&timeline->lock);
We prepared for the functions called inside the loop earlier, but the problem is the loop itself.
list_for_each_entry_safe will repeat the loop until next points back to temp.
struct kgsl_timeline_fence {
struct dma_fence base;
struct kgsl_timeline *timeline;
struct list_head node;
};
Because of automatic variable initialization, all fields including the node field are initialized, so we must construct a valid list_head structure ourselves. This means the following conditions must be met.
- The
next pointer must point to a valid pointer, and to prevent a crash caused by the functions inside the loop, it must point to an appropriate fake object.
- Eventually, one of the
next pointers must point back to temp, which exists on the kernel stack.
The first condition can be solved via an ion buffer at a predictable address, but the second condition is quite tricky.
Let's return for a moment to the dma_fence_signal_locked function that was called inside the loop.
int dma_fence_signal_locked(struct dma_fence *fence)
{
...
if (unlikely(test_and_set_bit(DMA_FENCE_FLAG_SIGNALED_BIT, //<-- 1.
&fence->flags)))
return -EINVAL;
/* Stash the cb_list before replacing it with the timestamp */
list_replace(&fence->cb_list, &cb_list); //<-- 2.
...
list_for_each_entry_safe(cur, tmp, &cb_list, node) { //<-- 3.
INIT_LIST_HEAD(&cur->node);
cur->func(fence, cur);
}
return 0;
}
The function above will be called with each of the two fences stored in the temp list (the fence reallocated via sendmsg and the fence linked to the ion buffer) passed one at a time as an argument.
As mentioned earlier, entering the code at 3 and calling the cur→func function causes a crash, so this must be avoided.
Therefore, to avoid it, the cb_list list must be an empty list.
To do this, we need to know our own address, so it's hard to accomplish with the fence reallocated via sendmsg, but it is possible with the fence residing in the ion buffer at a predictable address.
Therefore, if we set both the next and prev pointers to point to fence.cb_list and then enter the code at 2, the list_replace function is called first.
static inline void list_replace(struct list_head *old,
struct list_head *new)
{
//old->next = &(fence->cb_list)
new->next = old->next;
//new->next = &(fence->cb_list) => fence->cb_list.prev = &cb_list
new->next->prev = new;
//new->prev = fence->cb_list.prev => &cb_list
new->prev = old->prev;
//&cb_list->next = &cb_list
new->prev->next = new;
}
Through the call of the function above, the address of the kernel stack variable cb_list is written into fence→cb_list.prev, a field that exists in the ion buffer.
Since the ion buffer is a memory region mapped into userspace via the mmap function that we can read from and write to at any time, we can leak the kernel stack address stored in that field and compute the address of the temp list.
Now that we have obtained the address of the temp list, if we write it into the next address of the fake object in the ion buffer, we satisfy the second condition as well and can escape the loop.
Freelist Hijacking
Now we have resolved all of the kernel crash problems caused by object reallocation.
Now it's time to use the aforementioned DFB primitive.
Since a DFB primitive exists, if we allocate two objects with the same handler for the same memory chunk and then free one, that object is freed and becomes a freed chunk, and its first 8 bytes become the fd (freelist) pointer.
In this state, by accessing the remaining object, we can overwrite that fd pointer and thereby manipulate the location of the next dynamic allocation. Of course, several conditions must be satisfied, but they aren't too difficult.
As the object used to manipulate the fd pointer, we use the signalfd object. This object allocates an 8-byte object to store a mask, and the mask value stored in that object is a value the user can control with small restrictions.
The object is also freed when the signalfd file is closed, so controlling its lifetime is convenient.
By setting the fd pointer to an address within the ion buffer region via signalfd, subsequent heap memory becomes freely readable and writable, which is very convenient for exploitation.
The main obstacle in this process is kfree_rcu.
As we saw earlier, until the temp address is written into the next pointer, the loop keeps iterating.
That is, even after dma_fence_put is called, invokes kfree_rcu, and dma_fence_put returns, the CPU on which kfree_rcu was called will still be iterating the loop, and because that CPU is busily iterating while holding the spinlock, kfree_rcu will free the chunk on a different CPU.
Knowing which CPU it was freed on is necessary to reallocate the freed chunk as an arbitrary object, so this greatly affects the stability of the exploit.
To solve this, we simply perform the heap spray with a slight delay on each CPU.
Also, we can improve exploit stability by allocating a signalfd object immediately after the sendmsg object used at the beginning of the exploit design is freed, so as not to lose control.
Device Memory Mirroring Attack
Kernel drivers sometimes need to map memory into userspace, and for this their structures may hold pointers to a page structure or an sg_table structure.
This applies to the ion driver as well: as we saw earlier, to map memory into userspace via mmap, the ion_buffer object holds an sg_table structure.
struct sg_table {
struct scatterlist *sgl; /* the list */
unsigned int nents; /* number of mapped entries */
unsigned int orig_nents; /* original size of list */
};
struct scatterlist {
unsigned long page_link;
unsigned int offset;
unsigned int length;
dma_addr_t dma_address;
#ifdef CONFIG_NEED_SG_DMA_LENGTH
unsigned int dma_length;
#endif
};
The page_link field is the encoded form of the page pointer that points to the ion_buffer structure's backing store.
int ion_heap_map_user(struct ion_heap *heap, struct ion_buffer *buffer,
struct vm_area_struct *vma)
{
struct sg_table *table = buffer->sg_table;
...
for_each_sg(table->sgl, sg, table->nents, i) {
struct page *page = sg_page(sg);
...
//Maps pages to user space
ret = remap_pfn_range(vma, addr, page_to_pfn(page), len,
vma->vm_page_prot);
...
}
return 0;
}
When the mmap function is called, the backing store pointed to by page_link is mapped into userspace as shown above.
The page pointer is the value of the page's physical address shifted, plus a fixed offset.
Therefore, by manipulating page_link in this structure, we can map a desired kernel page into userspace.
KASLR only randomizes the virtual addresses that are linked to fixed physical addresses, and on many devices the kernel image is mapped at a fixed physical address, so being able to manipulate page_link alone is sufficient to exploit.
However, on Samsung devices, KASLR is applied to physical addresses as well. (More precisely, the intermediate physical address that the kernel perceives is not the real physical address but a virtual address assigned by the hypervisor.)
Therefore, in this exploit targeting Samsung devices, a physical address leak is also needed.
struct ion_buffer {
struct list_head list;
struct ion_heap *heap;
...
};
Fortunately, ion_buffer has an ion_heap field used to allocate the ion_buffer's backing stores.
struct ion_heap {
struct plist_node node;
enum ion_heap_type type;
struct ion_heap_ops *ops;
...
}
The ion_heap field has an ops field that points to the vtable for that object.
Since that vtable is a global object of the kernel image, we can bypass KASLR through the following process.
struct ion_buffer {
struct list_head list;
struct ion_heap *heap;
unsigned long flags;
...
- Find an
ion_buffer.
Using the flags field makes this easy.
When allocating an ion_buffer object via the ION_IOC_ALLOC ioctl, we can pass 4 bytes of arbitrary data as flags as a magic value, giving each ion_buffer a unique ID.
Using this, we can find the ion_buffer allocated for the fake object.
- Once the
ion_buffer object is identified, read that object's heap pointer, which is the pointer to its ion_heap object. Since that pointer is an address in Kernel Low Memory, its offset from the physical address is a static, constant offset. Therefore, we can easily obtain the physical address of that address.
- Once we obtain the physical address of the
ion_heap object pointed to by the heap pointer, manipulate the ion_buffer's sg_table so that the backing store points to the page where ion_heap exists.
- Call the
mmap function with the ion_buffer's fd so that the page where ion_heap exists is mapped into userspace. Thus, the ops pointer that exists in ion_heap can be accessed directly from userspace.
As mentioned earlier, the ops pointer is the address of a global object in the kernel image, so through it we can obtain the base offset.
ion_buffer can also solve another problem.
void kfree(const void *x)
{
struct page *page;
void *object = (void *)x;
trace_kfree(_RET_IP_, x);
if (unlikely(ZERO_OR_NULL_PTR(x)))
return;
page = virt_to_head_page(x);
if (unlikely(!PageSlab(page))) { //<-------- check if the page allocated is a single page slab
unsigned int order = compound_order(page);
BUG_ON(!PageCompound(page)); //<-------- check if the page is allocated as part of a multipage slab
...
}
...
}
If the fake object's object is freed, kfree checks via PageSlab whether the page containing the object is a single page slab of the SLUB allocator.
If not, PageCompound checks whether the page is part of a larger slab.
Since these checks are performed on the page structure itself, which contains the page's metadata, as soon as the object is freed the checks will fail and a kernel crash will occur.
This problem can be solved by tampering with the page structure's metadata using the AAR (Arbitrary Address Read) and AAW (Arbitrary Address Write) primitives we currently have.
Here, the address of the page structure corresponding to a physical address can easily be obtained via a shift operation on the physical address and a conversion to a fixed offset.
But if we ensure that the fake object's object is never freed, we can solve it more easily.
Before the ion_buffer structure is freed, ion_buffer_destroy is called.
int ion_buffer_destroy(struct ion_device *dev, struct ion_buffer *buffer)
{
...
heap = buffer->heap;
...
if (heap->flags & ION_HEAP_FLAG_DEFER_FREE)
ion_heap_freelist_add(heap, buffer); //<--------- does not free immediately
else
ion_buffer_release(buffer);
return 0;
}
If the ion_heap structure has the ION_HEAP_FLAG_DEFER_FREE flag, the ion_buffer is not freed immediately. Instead, ion_heap_freelist_add is called and it is added to the ion_heap's free_list.
The ion_buffer object is freed later when needed, and only when the ION_HEAP_FLAG_DEFER_FREE flag is set.
Normally, ION_HEAP_FLAG_DEFER_FREE does not change during the lifetime of the ion_heap object.
However, using the AAW primitive, we can add ION_HEAP_FLAG_DEFER_FREE to ion_heap→flags, free the ion_buffer, and then remove ION_HEAP_FLAG_DEFER_FREE again, so that the ion_buffer object is only added to the free_list and never freed.
Furthermore, since the page containing the ion_heap object was mapped into userspace earlier to bypass KASLR, the flag tampering can be performed easily.
By spraying fake objects to fill them with ion_buffer and its dependent objects, we can ensure these objects are never freed and thus never cause a kernel crash.
SELinux Bypass
When SELinux is enabled, it can be in either permissive mode or enforcing mode.
In permissive mode, unauthorized accesses are only logged, not blocked.
SELinux's mode is set by the selinux_enforcing variable.
If that variable is 0, SELinux operates in permissive mode.
Normally, security-critical variables are protected by Samsung KDP (Kernel Data Protection). Variables are set to read-only via the __kdp_ro or __rkp_ro attribute. These attributes place those variables in a read-only page, and modifications to them are protected by hypervisor calls. However, on this kernel branch, that protection mechanism is not applied (???).
Therefore, we can simply overwrite selinux_enforcing with 0 to set SELinux to permissive mode.
There is a general method for bypassing SELinux, but this time there is a simpler method, so we use it.
Spawning Root Shell
On Samsung devices, the biggest obstacle to gaining root privileges is RKP (Realtime Kernel Protection). The most common way to gain root privileges on an Android device is to overwrite the cred-related objects of the process performing the exploit.
However, RKP protects against tampering with these, so this time we use a different method.
kworker calls the functions of the workqueue with root privileges.
Since we have an AAW primitive, if we insert an arbitrary function into that workqueue, kworker will call it with root privileges.
However, since kCFI is applied, the functions we can call are limited.
Fortunately, the call_usermodehelper_exec_work function can be called by kworker and can execute shell commands, so by using it we can execute arbitrary commands with root privileges.
By manipulating system_unbound_wq to add an entry containing the call_usermodehelper_exec_work function pointer, we can bypass all protection mechanisms and spawn a root shell.