I am looking for simple way to split parenthesized lists that come out of IMAP responses into Python lists or tuples. I want to go from
\'(BODYSTRUCTURE (\"t
Taking out only internal part of the server answer containing actualy the body structure:
struct = ('(((("TEXT" "PLAIN" ("CHARSET" "ISO-8859-1") NIL NIL "7BIT" 16 2)'
'("TEXT" "HTML" ("CHARSET" "ISO-8859-1") NIL NIL "QUOTED-PRINTABLE"'
' 392 6) "ALTERNATIVE")("IMAGE" "GIF" ("NAME" "538.gif") '
'"<538@goomoji.gmail>" NIL "BASE64" 172)("IMAGE" "PNG" ("NAME" '
'"4F4.png") "" NIL "BASE64" 754) "RELATED")'
'("IMAGE" "JPEG" ("NAME" "avatar_airbender.jpg") NIL NIL "BASE64"'
' 157924) "MIXED")')
Next step is to replace some tokens, what would prepair string to transform into python types:
struct = struct.replace(' ', ',').replace(')(', '),(')
Using built-in module compiler to parse our structure:
import compiler
expr = compiler.parse(struct.replace(' ', ',').replace(')(', '),('), 'eval')
Performing simple recursive function to transform expression:
def transform(expression):
if isinstance(expression, compiler.transformer.Expression):
return transform(expression.node)
elif isinstance(expression, compiler.transformer.Tuple):
return tuple(transform(item) for item in expression.nodes)
elif isinstance(expression, compiler.transformer.Const):
return expression.value
elif isinstance(expression, compiler.transformer.Name):
return None if expression.name == 'NIL' else expression.name
And finally we get the desired result as nested python tuples:
result = transform(expr)
print result
(((('TEXT', 'PLAIN', ('CHARSET', 'ISO-8859-1'), None, None, '7BIT', 16, 2), ('TEXT', 'HTML', ('CHARSET', 'ISO-8859-1'), None, None, 'QUOTED-PRINTABLE', 392, 6), 'ALTERNATIVE'), ('IMAGE', 'GIF', ('NAME', '538.gif'), '<538@goomoji.gmail>', None, 'BASE64', 172), ('IMAGE', 'PNG', ('NAME', '4F4.png'), '', None, 'BASE64', 754), 'RELATED'), ('IMAGE', 'JPEG', ('NAME', 'avatar_airbender.jpg'), None, None, 'BASE64', 157924), 'MIXED')
From where we can recognize different headers of body structure:
text, attachments = (result[0], result[1:])